New England Section of the American Urological Association

NEAUA Home NEAUA Home Past & Future Meetings Past & Future Meetings

Back to 2026 Abstracts


Constraining Large Language Models to Clinical Algorithms for Safer, Evidence-Based Clinical Decision Support
Ashti Shah, MD1, Ajay Patel, PhD2, Marianne Casilla-Lennon, MD1, Michael Leapman, MD1, Joshua Sterling, MD1.
1Yale University, New Haven, CT, USA, 2University of Pennsylvania, Philadelphia, PA, USA.


Introduction: Adhering to clinical algorithms improves outcomes but is cognitively demanding. While LLMs offer decision support, they risk hallucinations and inaccurate recommendations. We hypothesize that a "Guideline Mode," constraining LLMs to structured pathways, enhances safety and physician trust.
Methods: Published evidence-based clinical algorithm PDFs for urologic conditions were converted into standardized JSON graph representations preserving decision tree structure. With patient data as input, the Urology Copilot chat application uses retrieval-augmented-generation (RAG) to find the relevant algorithm and constrains LLM recommendations to strictly follow the pathway. The application identifies the current step in the algorithm from the HPI and provides evidence quotations supporting the selection for auditability. Performance was evaluated on synthetic HPIs by assessing: (1) correct algorithm selection, (2) recommendation of appropriate next steps, and (3) ability to request additional information. Outcomes were compared to standard GPT-5 outputs. Thirty synthetic HPIs were created for 10 urologic diseases. 
Results: Urology Copilot correctly identified the clinical algorithm 93% of the time versus 60% for standard GPT-5 (p = 0.002). When selecting the correct algorithm, Copilot recommended the next step 100% of the time and requested more information when needed. Sixteen percent of HPIs required more than one algorithm; Copilot correctly selected the second algorithm and identified the correct step every time. Of standard GPT-5 outputs, 87% were clinically reasonable, though 30% were guideline-inconsistent. 
Conclusion: Confining LLM outputs to structured, guideline-based pathways improves accuracy, safety, and auditability. Urology Copilot outperformed standard LLMs in selecting algorithms and recommending next steps, demonstrating the potential of guideline-constrained models to enhance physician adherence while reducing guideline-inconsistent recommendations. Work is underway to evaluate Copilot on HIPAA-protected patient HPIs.  
Back to 2026 Abstracts