Patient records, literature, and now AI-generated summaries have grown past what any one person can read in full. Medical Thinking is about the flow of information in medicine — what is the most pertinent information for a given reader, at a given moment. That question starts with summarization, done rigorously enough to trust.
Work toward making the summarization and flow of medical information evaluable, not just fluent.
A two-layer framework — universal criteria plus specialty-specific salience — for grading AI-generated clinical summaries against consensus-ratified criteria rather than reference summaries.
What a cardiologist needs before clinic differs from what primary care needs at an annual exam. Tailoring has to scale past per-user idiosyncrasy.
A summary that drops a critical lab trend still reads as complete. Evaluating only what's said, never what's missing, misses the most dangerous failure mode.
No single "gold" summary generalizes. What scales is a validated, consensus-ratified rubric — and a validated process for applying it.