Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction
Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state. We propose CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification. CAST uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations, labels SAE latents with an LLM-assisted interpretation pipeline and ICD-10 retrieval constraints, suppresses verified artifact latents
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%Addressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming - Nature →
- LinkedLinked via arxiv author · 85%Jin Mu →
“Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction”
- LinkedLinked via arxiv author · 85%Guanhua Chen →
“Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction”
