From Benchmarks to Courtrooms: A Capability–Reliability–Accountability Survey of Large Language Models in Legal Practice

Authors

  • 钰颖 张 吉利学院
  • 馨宇 宋

Keywords:

hallucination;, retrieval-augmented generation;, judgment prediction;, neuro-symbolic reasoning;, evaluation validity

Abstract

Large language models have entered legal practice rapidly, yet systems that score highly on legal benchmarks have fabricated judicial citations in filed briefs—a contradiction that resource-inventory surveys do not explain. This survey reviews 25 core sources (2019–2026, plus the 1987–2003 symbolic genealogy) through an original capability–reliability–accountability (CRA) framework; reported figures are quoted only from peer-reviewed studies, cited case law, or first-party technical reports. Measured capability has matured across three benchmark generations, with frontier models exceeding 85% on aggregate legal-reasoning tasks. Reliability has not kept pace: benchmark validity limits, citation hallucination rates of 17–43% even under retrieval grounding, and pipeline-level retrieval failures independently sever the link between scores and trustworthiness. Accountability responses—high-risk regulation, neuro-symbolic scaffolding descending from the HYPO tradition, and rubric-bound LLM-as-a-judge evaluation—converge on a single strategy: re-embedding statistical systems in inspectable structure. The CRA framework explains the benchmark-to-courtroom gap, positions the field's open problems, and transfers to other high-risk professional domains.

References

Harvey. (2025). BigLaw Bench: A deep dive on retrieval. Harvey AI. https://www.harvey.ai/blog/biglaw-bench-retrieval

Mata v. Avianca, Inc., No. 22-cv-1461 (S.D.N.Y. 2023).

Guha, N., Ho, D. E., & Nyarko, J. (2023). LegalBench: A collaboratively built benchmark for measuring legal reasoning of large language models. In Advances in Neural Information Processing Systems 36 (Datasets and Benchmarks Track).

Hou, Z., Ye, Z., Zeng, N., Hao, T., & Zeng, K. (2025). Large language models meet legal artificial intelligence: A survey. arXiv preprint arXiv:2509.09969.

Dehghani, F. (2025). Large language models in legal systems: A survey. Humanities and Social Sciences Communications, 12. https://www.nature.com/articles/s41599-025-05924-3

Chalkidis, I., Fergadiotis, M., Malakasiotis, P., Aletras, N., & Androutsopoulos, I. (2020). LEGAL-BERT: The Muppets straight out of Law School. In Findings of the Association for Computational Linguistics: EMNLP 2020.

Tagarelli, A., & Simeri, A. (2023). Italian-legal-bert models for improving natural language processing in the Italian legal domain. Computer Law & Security Review, 48, 105787.

Chalkidis, I., Fergadiotis, M., Malakasiotis, P., & Androutsopoulos, I. (2019). Neural legal judgment prediction in English. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.

Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems 35.

Chalkidis, I., Jana, A., Hartung, D., Bommarito, M., Androutsopoulos, I., Katz, D. M., & Aletras, N. (2022). LexGLUE: A benchmark dataset for legal language understanding in English. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics.

Fei, Z., Shen, X., Zhu, D., Zhou, F., Han, Z., Zhang, S., Chen, K., Zhang, Z., Xu, E., Lin, B. Y., Geiger, A., Yu, T., & Chen, H. (2023). LawBench: Benchmarking legal knowledge of large language models. arXiv preprint arXiv:2309.16289.

Choi, J. H., Hickman, K. E., Monahan, A. B., & Schwarcz, D. B. (2023). ChatGPT goes to law school. Journal of Legal Education, 71(3), 387.

Bommarito, M. J., II, Katz, D. M., & Detterman, A. (2022). GPT takes the bar exam. arXiv preprint arXiv:2112.01821.

OpenAI. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.

Hu, Y., Yu, Y., Gan, L., Wei, B., Kuang, K., & Wu, F. (2025). Evaluating test-time scaling LLMs for legal reasoning: OpenAI o1, DeepSeek-R1, and beyond. In Findings of the Association for Computational Linguistics: EMNLP 2025.

Medvedeva, M., Wieling, M., & Vols, M. (2023). Rethinking the field of automatic prediction of court decisions. Artificial Intelligence and Law, 31(2), 195-212.

Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal of Legal Analysis, 16(1), 64-93.

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems 33.

European Parliament and Council. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union.

Rissland, E. L., & Ashley, K. D. (1987). HYPO: A case-based system for trade secrets law. In Proceedings of the First International Conference on Artificial Intelligence and Law (ICAIL '87). ACM.

Bruninghaus, L., & Ashley, K. D. (2003). Predicting outcomes of case-based legal arguments. In Proceedings of the 9th International Conference on Artificial Intelligence and Law (ICAIL '03). ACM.

Bench-Capon, T. J. M. (2003). Persuasion in practical argument using value-based argumentation frameworks. Journal of Logic and Computation, 13(3), 429-448.

Chen, L., Cai, Y., Hou, Z., & Dong, J. S. (2025). Towards trustworthy legal AI through LLM agents and formal reasoning. arXiv preprint arXiv:2511.21033.

Pradhan, A., Ortan, A., Verma, A., & Seshadri, M. (2025). LLM-as-a-judge: Rapid evaluation of legal document recommendation for retrieval-augmented generation. arXiv preprint arXiv:2509.12382.

Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems 36 (Datasets and Benchmarks Track).

Downloads

Published

2026-09-01

How to Cite

张钰., & 宋馨. (2026). From Benchmarks to Courtrooms: A Capability–Reliability–Accountability Survey of Large Language Models in Legal Practice. International Journal of Advanced AI Applications, 2(9), 80–98. Retrieved from https://www.dawnclarity.press/index.php/ijaaa/article/view/186