IASCI Research Publishing
Translational Medicine and Digital Health

Surface Robustness in Medical Language Agents

Read & download PDF
Abstract

How should a medical agent be tested when semantically equivalent requests produce different actions? This review answers by treating surface-form stress testing as a property of a sociotechnical workflow rather than a feature that can be read from average accuracy. The focal setting is clinical decision support under adversarial and accidental wording variation, where an error can be fluent, repeatable, and clinically consequential even when aggregate benchmark accuracy is high. Evidence from the assigned publications is synthesized with foundational studies of calibration, distribution shift, causal structure, and responsible deployment. Four requirements follow: preserve the lineage of patient signals, clinical language, retrieved evidence, attack variants, and workflow metadata; measure stability across relevant perturbations; connect confidence to a specific action; and maintain a route for human challenge and correction. The framework distinguishes descriptive performance from decision utility and separates uncertainty about the world from uncertainty created by the model and its evaluator. It also shows why faster inference or richer reasoning is valuable only when it improves a defined decision under a transparent resource budget. The article is a literature review and research agenda, not a report of a newly completed trial.

Keywords
surface robustnessmedical language agentsdecisionmedicallanguageaccuracyclinical
References
  1. Hu, Saisai. "Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks." *arXiv preprint arXiv:2605.08257* (2026).
  2. Li, Yuanhao, et al. "DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs." *Proceedings of the AAAI Conference on Artificial Intelligence* 40.35 (2026): 29530-29537.
  3. Sang, Yinghao. "Adaptive Quantization Strategies for Robust ML Inference Under Distribution Shift." *Proceedings of the 2026 5th International Conference on Cyber Security, Artificial Intelligence and Digital Economy* (2026): 362-368.
  4. Su, Tongli, et al. "Cross-Modal Generative Framework for Signal Translation from Fetal-Maternal Electrocardiograms to Fetal Doppler Waveforms." *arXiv preprint arXiv:2607.08073* (2026).
  5. Zhou, Yongxi, et al. "Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity." *arXiv preprint arXiv:2608.02665* (2026).
  6. Geifman, Yonatan, and Ran El-Yaniv. "Selective Classification for Deep Neural Networks." *Advances in Neural Information Processing Systems*, vol. 30, 2017.
  7. Ribeiro, Marco Tulio, Sameer Singh, and Carlos Guestrin. "Why Should I Trust You?: Explaining the Predictions of Any Classifier." *Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining*, 2016, pp. 1135-1144.
  8. Singhal, Karan, et al. "Large Language Models Encode Clinical Knowledge." *Nature*, vol. 620, 2023, pp. 172-180.
  9. Wiens, Jenna, et al. "Do No Harm: A Roadmap for Responsible Machine Learning for Health Care." *Nature Medicine*, vol. 25, 2019, pp. 1337-1340.
  10. Rajkomar, Alvin, Jeffrey Dean, and Isaac Kohane. "Machine Learning in Medicine." *New England Journal of Medicine*, vol. 380, 2019, pp. 1347-1358.
  11. World Health Organization. *Ethics and Governance of Artificial Intelligence for Health*. World Health Organization, 2021.
  12. National Institute of Standards and Technology. *Artificial Intelligence Risk Management Framework (AI RMF 1.0)*. U.S. Department of Commerce, 2023.
  13. Sculley, D., et al. "Hidden Technical Debt in Machine Learning Systems." *Advances in Neural Information Processing Systems*, vol. 28, 2015.
  14. Hendrycks, Dan, and Thomas Dietterich. "Benchmarking Neural Network Robustness to Common Corruptions and Perturbations." *International Conference on Learning Representations*, 2019.
  15. Goodfellow, Ian J., Jonathon Shlens, and Christian Szegedy. "Explaining and Harnessing Adversarial Examples." *International Conference on Learning Representations*, 2015.
Publication details
Journal
Translational Medicine and Digital Health
Volume
1 (2026)
Issue
1 ยท Forthcoming issue
Article number
tmdh20260001
License
CC BY 4.0