Benchmarking Autonomy: A Taxonomy of Human-in-the-Loop Checkpoints for AI-Generated Transformation Code
DOI:
https://doi.org/10.22399/ijcesen.5462Keywords:
AI-generated code, data transformation, SQL, Apache Spark, dbt, data qualityAbstract
Generative code systems can produce executable SQL, Spark, and dbt transformations with substantially less manual effort, but syntactic validity does not establish semantic correctness. A transformation may compile, complete successfully, and preserve an expected schema while silently changing the meaning of downstream data. This article presents a controlled pre-production benchmark and a taxonomy of human-in-the-loop checkpoints for governing that risk. The benchmark contains 12 realistic transformation tasks spanning relational SQL, distributed Spark processing, and dbt-style analytical modeling. Nine candidates contain seeded defects and three are correct controls. Three automated checkpoint families are evaluated independently and in combination: structural schema checks, declarative data-quality rules, and differential comparison against a trusted baseline. Schema checks and data-quality rules each detect one of nine defects (11.1%), and their detections overlap. Trusted-baseline comparison detects seven defects (77.8%). The union of all automated checkpoints detects eight defects (88.9%), leaving one low-magnitude semantic error that preserves schema, distributions, row counts, and tolerance-bounded aggregate values. The residual defect is identified only when the transformation logic and business intent are reviewed together. The findings support three conclusions. First, validation depth matters more than check quantity: semantic oracles materially outperform purely structural controls. Second, checkpoint portfolios exhibit diminishing returns when multiple rules target the same visible symptom. Third, human review should not be treated as an undifferentiated manual gate; it should be targeted toward transformations with monetary, temporal, incremental, rare-category, or tolerance-sensitive semantics. The article contributes a five-level checkpoint taxonomy, a reproducible defect taxonomy, task-level detection evidence, and a risk-based escalation policy suitable for pre-production DataOps and analytics engineering workflows. The analysis is artifact-centric and does not rank named code-generation models because generation logs and repeated model samples were not part of the benchmark. Because the benchmark is intentionally small, the numerical estimates are reported with uncertainty and are not generalized as population-level industry rates.
References
[1] R. Y. Wang and D. M. Strong, “Beyond accuracy: What data quality means to data consumers,” J. Manage. Inf. Syst., vol. 12, no. 4, pp. 5–33, 1996, doi: 10.1080/07421222.1996.11518099. DOI: https://doi.org/10.1080/07421222.1996.11518099
[2] L. L. Pipino, Y. W. Lee, and R. Y. Wang, “Data quality assessment,” Commun. ACM, vol. 45, no. 4, pp. 211–218, 2002, doi: 10.1145/505248.506010. DOI: https://doi.org/10.1145/505248.506010
[3] C. Batini, C. Cappiello, C. Francalanci, and A. Maurino, “Methodologies for data quality assessment and improvement,” ACM Comput. Surv., vol. 41, no. 3, Art. no. 16, 2009, doi: 10.1145/1541880.1541883. DOI: https://doi.org/10.1145/1541880.1541883
[4] D. M. Strong, Y. W. Lee, and R. Y. Wang, “Data quality in context,” Commun. ACM, vol. 40, no. 5, pp. 103–110, 1997, doi: 10.1145/253769.253804. DOI: https://doi.org/10.1145/253769.253804
[5] D. Sculley et al., “Hidden technical debt in machine learning systems,” in Advances in Neural Information Processing Systems 28, 2015, pp. 2503–2511, doi: 10.5555/2969442.2969519.
[6] E. Breck, S. Cai, E. Nielsen, M. Salib, and D. Sculley, “The ML test score: A rubric for ML production readiness and technical debt reduction,” in Proc. IEEE Int. Conf. Big Data, 2017, pp. 1123–1132, doi: 10.1109/BigData.2017.8258038. DOI: https://doi.org/10.1109/BigData.2017.8258038
[7] D. Baylor et al., “TFX: A TensorFlow-based production-scale machine learning platform,” in Proc. 23rd ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, 2017, pp. 1387–1395, doi: 10.1145/3097983.3098021. DOI: https://doi.org/10.1145/3097983.3098021
[8] S. Schelter et al., “Automating large-scale data quality verification,” Proc. VLDB Endow., vol. 11, no. 12, pp. 1781–1794, 2018, doi: 10.14778/3229863.3229867. DOI: https://doi.org/10.14778/3229863.3229867
[9] S. Schelter et al., “Unit testing data with Deequ,” in Proc. 2019 Int. Conf. Manage. Data, 2019, pp. 1993–1996, doi: 10.1145/3299869.3320210. DOI: https://doi.org/10.1145/3299869.3320210
[10] L. E. Lwakatare, E. Range, I. Crnkovic, and J. Bosch, “On the experiences of adopting automated data validation in an industrial machine learning project,” in Proc. IEEE/ACM ICSE-SEIP, 2021, pp. 248–257, doi: 10.1109/ICSE-SEIP52600.2021.00034. DOI: https://doi.org/10.1109/ICSE-SEIP52600.2021.00034
[11] E. Caveness et al., “TensorFlow Data Validation: Data analysis and validation in continuous ML pipelines,” in Proc. 2020 ACM SIGMOD Int. Conf. Manage. Data, 2020, pp. 2793–2796, doi: 10.1145/3318464.3384707. DOI: https://doi.org/10.1145/3318464.3384707
[12] J. Song and Y. He, “Auto-Validate: Unsupervised data validation using data-domain patterns inferred from data lakes,” in Proc. 2021 ACM SIGMOD Int. Conf. Manage. Data, 2021, pp. 1678–1691, doi: 10.1145/3448016.3457250. DOI: https://doi.org/10.1145/3448016.3457250
[13] E. T. Barr, M. Harman, P. McMinn, M. Shahbaz, and S. Yoo, “The oracle problem in software testing: A survey,” IEEE Trans. Softw. Eng., vol. 41, no. 5, pp. 507–525, 2015, doi: 10.1109/TSE.2014.2372785. DOI: https://doi.org/10.1109/TSE.2014.2372785
[14] Y. Jia and M. Harman, “An analysis and survey of the development of mutation testing,” IEEE Trans. Softw. Eng., vol. 37, no. 5, pp. 649–678, 2011, doi: 10.1109/TSE.2010.62. DOI: https://doi.org/10.1109/TSE.2010.62
[15] R. Just, D. Jalali, and M. D. Ernst, “Defects4J: A database of existing faults to enable controlled testing studies for Java programs,” in Proc. Int. Symp. Softw. Testing Anal., 2014, pp. 437–440, doi: 10.1145/2610384.2628055. DOI: https://doi.org/10.1145/2610384.2628055
[16] T. Y. Chen et al., “Metamorphic testing: A review of challenges and opportunities,” ACM Comput. Surv., vol. 51, no. 1, Art. no. 4, 2018, doi: 10.1145/3143561. DOI: https://doi.org/10.1145/3143561
[17] M. Rigger and Z. Su, “Testing database engines via pivoted query synthesis,” in Proc. 14th USENIX Symp. Oper. Syst. Design Implement., 2020, pp. 667–682, doi: 10.48550/arXiv.2001.04174.
[18] M. Rigger and Z. Su, “Detecting optimization bugs in database engines via non-optimizing reference engine construction,” in Proc. 28th ACM ESEC/FSE, 2020, pp. 1140–1152, doi: 10.1145/3368089.3409710. DOI: https://doi.org/10.1145/3368089.3409710
[19] M. Rigger and Z. Su, “Finding bugs in database systems via query partitioning,” Proc. ACM Program. Lang., vol. 4, OOPSLA, Art. no. 211, 2020, doi: 10.1145/3428279. DOI: https://doi.org/10.1145/3428279
[20] R. Zhong et al., “SQUIRREL: Testing database management systems with language validity and coverage feedback,” in Proc. ACM CCS, 2020, pp. 955–967, doi: 10.1145/3372297.3417260. DOI: https://doi.org/10.1145/3372297.3417260
[21] V. S. P. Dintyala, A. Narechania, and J. Arulraj, “SQLCheck: Automated detection and diagnosis of SQL anti-patterns,” in Proc. 2020 ACM SIGMOD, 2020, pp. 2331–2345, doi: 10.1145/3318464.3389754. DOI: https://doi.org/10.1145/3318464.3389754
[22] Q. Zhang, J. Wang, M. A. Gulzar, R. Padhye, and M. Kim, “BigFuzz: Efficient fuzz testing for data analytics using framework abstraction,” in Proc. 35th IEEE/ACM ASE, 2020, doi: 10.1145/3324884.3416641. DOI: https://doi.org/10.1145/3324884.3416641
[23] K. Kallas, F. Niksic, C. Stanford, and R. Alur, “DiffStream: Differential output testing for stream processing programs,” Proc. ACM Program. Lang., vol. 4, OOPSLA, Art. no. 153, 2020, doi: 10.1145/3428221. DOI: https://doi.org/10.1145/3428221
[24] M. A. Gulzar, S. Mardani, M. Musuvathi, and M. Kim, “White-box testing of big data analytics with complex user-defined functions,” in Proc. 27th ACM ESEC/FSE, 2019, pp. 290–301, doi: 10.1145/3338906.3338953. DOI: https://doi.org/10.1145/3338906.3338953
[25] Z. Feng et al., “CodeBERT: A pre-trained model for programming and natural languages,” in Findings of EMNLP, 2020, pp. 1536–1547, doi: 10.18653/v1/2020.findings-emnlp.139. DOI: https://doi.org/10.18653/v1/2020.findings-emnlp.139
[26] D. Guo et al., “GraphCodeBERT: Pre-training code representations with data flow,” in Proc. ICLR, 2021, doi: 10.48550/arXiv.2009.08366.
[27] W. U. Ahmad, S. Chakraborty, B. Ray, and K.-W. Chang, “Unified pre-training for program understanding and generation,” in Proc. NAACL-HLT, 2021, pp. 2655–2668, doi: 10.18653/v1/2021.naacl-main.211. DOI: https://doi.org/10.18653/v1/2021.naacl-main.211
[28] M. Chen et al., “Evaluating large language models trained on code,” 2021, doi: 10.48550/arXiv.2107.03374.
[29] J. Austin et al., “Program synthesis with large language models,” 2021, doi: 10.48550/arXiv.2108.07732.
[30] J. A. Prenner and R. Robbes, “Automatic program repair with OpenAI Codex: Evaluating QuixBugs,” 2021, doi: 10.48550/arXiv.2111.03922.
[31] S. Amershi et al., “Guidelines for human-AI interaction,” in Proc. 2019 CHI Conf. Human Factors Comput. Syst., 2019, pp. 1–13, doi: 10.1145/3290605.3300233. DOI: https://doi.org/10.1145/3290605.3300233
[32] X. Wu, L. Xiao, Y. Sun, J. Zhang, T. Ma, and L. He, “A survey of human-in-the-loop for machine learning,” 2021, doi: 10.48550/arXiv.2108.00941.
[33] A. Crisan and B. Fiore-Gartland, “Fits and starts: Enterprise use of AutoML and the role of humans in the loop,” in Proc. 2021 CHI, 2021, doi: 10.1145/3411764.3445775. DOI: https://doi.org/10.1145/3411764.3445775
[34] N. Sambasivan et al., “Everyone wants to do the model work, not the data work: Data cascades in high-stakes AI,” in Proc. 2021 CHI, 2021, doi: 10.1145/3411764.3445518. DOI: https://doi.org/10.1145/3411764.3445518
[35] C. Wohlin et al., Experimentation in Software Engineering. Berlin, Germany: Springer, 2012, doi: 10.1007/978-3-642-29044-2. DOI: https://doi.org/10.1007/978-3-642-29044-2
[36] B. A. Kitchenham et al., “Preliminary guidelines for empirical research in software engineering,” IEEE Trans. Softw. Eng., vol. 28, no. 8, pp. 721–734, 2002, doi: 10.1109/TSE.2002.1027796. DOI: https://doi.org/10.1109/TSE.2002.1027796
[37] B. Fitzgerald and K.-J. Stol, “Continuous software engineering: A roadmap and agenda,” J. Syst. Softw., vol. 123, pp. 176–189, 2017, doi: 10.1016/j.jss.2015.06.063. DOI: https://doi.org/10.1016/j.jss.2015.06.063
[38] R. Gollapudi, “Risk-controlled near-zero-downtime Oracle database migration using GoldenGate,” Int. J. Comput. Exp. Sci. Eng., vol. 8, no. 3, pp. 113–123, 2022, doi: 10.22399/ijcesen.5382. DOI: https://doi.org/10.22399/ijcesen.5382
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2022 International Journal of Computational and Experimental Science and Engineering

This work is licensed under a Creative Commons Attribution 4.0 International License.