Can LLM-Generated Backend Services Meet Performance Slos? A Controlled Study of Explicit Requirements and Measured-Feedback Repair

  • Authors

    • Priyank Agrawal Master of Science in Information Systems, Northeastern University: Boston, Massachusetts, US. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-V9I3P103

    Published 08-12-2026

  • Large Language Models, Coding Agents, Software Performance, Service-Level Objectives, Backend Services, Execution Feedback, Benchmarking, Reproducibility

    Issue

    Section

    Articles

    How to Cite

    [1]
    P. Agrawal, “Can LLM-Generated Backend Services Meet Performance Slos? A Controlled Study of Explicit Requirements and Measured-Feedback Repair”, IJDEIC, vol. 9, no. 3, pp. 47–63, Aug. 2026, doi: 10.67228/30715717/IJDEIC-V9I3P103.
  • Abstract

    Production backend services must satisfy functional contracts and operational objectives simultaneously, yet most evaluations of code-generating language models still emphasize compilation, unit-test success, or repository issue resolution. This study asks whether a coding-agent workflow can generate small backend services that satisfy explicit service-level objectives (SLOs), and whether measured runtime feedback changes the result. We built SLOBench, a reproducible local harness that combines versioned prompts, deterministic functional validation, process-isolated FastAPI execution, concurrent HTTP load, CPU and memory sampling, provenance hashes, and implementation-level analysis. The controlled pilot evaluated three services, metadata lookup, a bounded cache-backed API, and concurrent aggregation, under three conditions: functional-only prompting, functional prompting plus an explicit SLO, and one-step repair after measured feedback. Five generation attempts per task produced 45 implementations; each implementation was measured three times, yielding 135 final runs. All implementations passed the functional oracle and every final load run had zero observed request errors. Adding SLO language alone reduced P95 latency by a median 2.2% across 15 paired generation-task comparisons (bootstrap 95% interval -2.7% to 10.7%; 9/15 improved). Measured-feedback repair improved all 15 pairs, with a median 16.8% reduction relative to the SLO-prompted version (95% interval 15.7% to 20.2%). The two strictest latency targets were nevertheless never reached. Source review further showed that all repaired implementations bypassed portions of the ordinary framework path and three hard-coded the disclosed hot request. Measured feedback therefore improved performance on the declared workload consistently, but the evidence does not establish production readiness. The findings motivate hidden acceptance workloads, calibrated baselines, stronger post-repair contract testing, and multi-model replication in performance-oriented coding-agent evaluation.

  • References

    [1] Bommasani, R., Liang, P., & Lee, T. (2023). Holistic evaluation of language models. Annals of the New York Academy of Sciences, 1525(1), 140–146. https://doi.org/10.1111/nyas.15007

    [2] Chen, B., Zhang, F., Nguyen, A., Zan, D., Lin, Z., Lou, J.-G., & Chen, W. (2023). CodeT: Code generation with generated tests. International Conference on Learning Representations. https://openreview.net/forum?id=ktrw68Cmu9c

    [3] Curtsinger, C., & Berger, E. D. (2013). STABILIZER: Statistically sound performance evaluation. In Proceedings of the 18th International Conference on Architectural Support for Programming Languages and Operating Systems (pp. 219–228). Association for Computing Machinery. https://doi.org/10.1145/2451116.2451141

    [4] Dean, J., & Barroso, L. A. (2013). The tail at scale. Communications of the ACM, 56(2), 74–80. https://doi.org/10.1145/2408776.2408794

    [5] Du, M., Tuan, L. A., Ji, B., Liu, Q., & Ng, S.-K. (2024). Mercury: A code efficiency benchmark for code large language models. Advances in Neural Information Processing Systems, 37, 16601–16622. https://doi.org/10.52202/079017-0529

    [6] Fan, A., Gokkaya, B., Harman, M., Lyubarskiy, M., Sengupta, S., Yoo, S., & Zhang, J. M. (2023). Large language models for software engineering: Survey and open problems. In 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE) (pp. 31–53). IEEE. https://doi.org/10.1109/ICSE-FoSE59343.2023.00008

    [7] Gan, Y., Zhang, Y., Cheng, D., Shetty, A., Rathi, P., Katarki, N., Bruno, A., Hu, J., Ritchken, B., Jackson, B., Hu, K., Pancholi, M., He, Y., Clancy, B., Colen, C., Wen, F., Leung, C., Wang, S., Zaruvinsky, L., ... Delimitrou, C. (2019a). An open-source benchmark suite for microservices and their hardware-software implications for cloud and edge systems. In Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems (pp. 3–18). Association for Computing Machinery. https://doi.org/10.1145/3297858.3304013

    [8] Gan, Y., Zhang, Y., Hu, K., Cheng, D., He, Y., Pancholi, M., & Delimitrou, C. (2019b). Seer: Leveraging big data to navigate the complexity of performance debugging in cloud microservices. In Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems (pp. 19–33). Association for Computing Machinery. https://doi.org/10.1145/3297858.3304004

    [9] Georges, A., Buytaert, D., & Eeckhout, L. (2007). Statistically rigorous Java performance evaluation. In Proceedings of the 22nd Annual ACM SIGPLAN Conference on Object-Oriented Programming Systems and Applications (pp. 57–76). Association for Computing Machinery. https://doi.org/10.1145/1297027.1297033

    [10] Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., & Steinhardt, J. (2021). Measuring coding challenge competence with APPS. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 1.

    [11] Huang, D., Qing, Y., Shang, W., Cui, H., & Zhang, J. M. (2024). EffiBench: Benchmarking the efficiency of automatically generated code. Advances in Neural Information Processing Systems, 37. https://doi.org/10.52202/079017-0367

    [12] Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., & Stoica, I. (2025). LiveCodeBench: Holistic and contamination free evaluation of large language models for code. International Conference on Learning Representations. https://proceedings.iclr.cc/paper_files/paper/2025/hash/94074dd5a072d28ff75a76dabed43767-Abstract-Conference.html

    [13] Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. R. (2024). SWE-bench: Can language models resolve real-world GitHub issues? International Conference on Learning Representations. https://openreview.net/forum?id=VTF8yNQM66

    [14] Kalibera, T., & Jones, R. E. (2013). Rigorous benchmarking in reasonable time. In Proceedings of the 2013 International Symposium on Memory Management (pp. 63–74). Association for Computing Machinery. https://doi.org/10.1145/2464157.2464160

    [15] Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). CodeRL: Mastering code generation through pretrained models and deep reinforcement learning. Advances in Neural Information Processing Systems, 35, 21314–21328. https://doi.org/10.52202/068431-1549

    [16] Leitner, P., & Cito, J. (2016). Patterns in the chaos: A study of performance variation and predictability in public IaaS clouds. ACM Transactions on Internet Technology, 16(3), 1–23. https://doi.org/10.1145/2885497

    [17] Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., Hubert, T., Choy, P., de Masson d'Autume, C., Babuschkin, I., Chen, X., Huang, P.-S., Welbl, J., Gowal, S., Cherepanov, A., ... Vinyals, O. (2022). Competition-level code generation with AlphaCode. Science, 378(6624), 1092–1097. https://doi.org/10.1126/science.abq1158

    [18] Liu, J., Xia, C. S., Wang, Y., & Zhang, L. (2023). Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems, 36, 21558–21572.

    [19] Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., & Clark, P. (2023). Self-Refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36. https://doi.org/10.52202/075280-2019

    [20] Mytkowicz, T., Diwan, A., Hauswirth, M., & Sweeney, P. F. (2009). Producing wrong data without doing anything obviously wrong! In Proceedings of the 14th International Conference on Architectural Support for Programming Languages and Operating Systems (pp. 265–276). Association for Computing Machinery. https://doi.org/10.1145/1508244.1508275

    [21] Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., & Xiong, C. (2023). CodeGen: An open large language model for code with multi-turn program synthesis. International Conference on Learning Representations. https://openreview.net/forum?id=iaYcJKpY2B_

    [22] Peng, Y., Gotmare, A. D., Lyu, M., Xiong, C., Savarese, S., & Sahoo, D. (2025). PerfCodeGen: Improving performance of LLM generated code with execution feedback. In 2025 IEEE/ACM 2nd International Conference on AI Foundation Models and Software Engineering (FORGE) (pp. 1–13). IEEE. https://doi.org/10.1109/FORGE66646.2025.00008

    [23] Qiu, R., Zeng, W. W., Ezick, J., Lott, C., & Tong, H. (2025). How efficient is LLM-generated code? A rigorous and high-standard benchmark. International Conference on Learning Representations. https://proceedings.iclr.cc/paper_files/paper/2025/hash/06694da057cb15fef11542270a592627-Abstract-Conference.html

    [24] Rzadca, K., Findeisen, P., Swiderski, J., Zych, P., Broniek, P., Kusmierek, J., Nowak, P., Strack, B., Witusowski, P., Hand, S., & Wilkes, J. (2020). Autopilot: Workload autoscaling at Google. In Proceedings of the Fifteenth European Conference on Computer Systems (Article 16, pp. 1–16). Association for Computing Machinery. https://doi.org/10.1145/3342195.3387524

    [25] Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36.

    [26] Shypula, A., Madaan, A., Zeng, Y., Alon, U., Gardner, J., Yang, Y., Hashemi, M., Neubig, G., Ranganathan, P., Bastani, O., & Yazdanbakhsh, A. (2024). Learning performance-improving code edits. International Conference on Learning Representations. https://proceedings.iclr.cc/paper_files/paper/2024/hash/ce65173b994cf7c925c71b482ee14a8d-Abstract-Conference.html

    [27] Smith, A. M., Katz, D. S., Niemeyer, K. E., & FORCE11 Software Citation Working Group. (2016). Software citation principles. PeerJ Computer Science, 2, e86. https://doi.org/10.7717/peerj-cs.86

    [28] Waghjale, S., Veerendranath, V., Wang, Z. Z., & Fried, D. (2024). ECCO: Can we improve model-generated code efficiency without sacrificing functional correctness? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 15362–15376). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.859

    [29] Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). SWE-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems, 37. https://doi.org/10.52202/079017-1601

    [30] Zheng, T., Zhang, G., Shen, T., Liu, X., Lin, B. Y., Fu, J., Chen, W., & Yue, X. (2024). OpenCodeInterpreter: Integrating code generation with execution and refinement. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 12834–12859). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.762

    [31] Begimher, D., Leo, C., Huang, J., Gaw, P., & Zheng, B. (2026). SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents. arXiv preprint arXiv:2604.12040.

    [32] Leo, C., Dykyi, A., Cortegaca, D., Begimher, D., & Jha, P. (2026). ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping. arXiv preprint arXiv:2607.27528.

  • Downloads