Scientific Reports posted a clinical annotation pipeline on September 29 that does not need a vendor login. One German hospital, 1,100 emergency-department notes, nine locally deployed open-weight models, a frozen zero-shot prompt. Temperature 0.0 was the setting that did not fall apart. Vomiting hit F1 0.985. Diarrhea did not beat the old rule list.
A day later, a software-vulnerability paper made the same methodological complaint in another building. RAG systems that looked strong on proprietary models shrank once someone reran them on open weights. We already said MiniCPM5-2B’s 2.8-point average is not a server and that leaderboards are not the job. This week is what a hospital does instead of a leaderboard.
The setup is a pipeline, not a bake-off
The authors wanted a reproducible way to pick a local model for binary symptom labels in surgical free text. Nausea, vomiting, diarrhea, dysuria. Human annotators labeled 1,100 German ED reports. One hundred notes went to prompt drafting and temperature tests. For each symptom, the remaining thousand were split 250 / 750, stratified, fixed seed. Nine models were screened on the 250. The winner went to the 750 against a negation-aware rule baseline.
They kept the prompt simple on purpose. No prompt Olympics. No claim that they found the best model in the world. The question was whether a screening workflow can choose a local model that is good enough to label notes you cannot send to a cloud.
That is a different genre from open weights taking tokens while closed models keep the bill. Nobody here is selling an API. The models have names you can actually pull: gemma2:27b for nausea, mistral-large:123b for vomiting, llama3.3:70b for diarrhea and dysuria. Those were the screening winners, not a committee’s favorite brand.
Median inference on the chosen models sat between 0.298 and 1.653 seconds per report. That is the operational sentence hiding under the F1 table. A pipeline that needs a minute per note is a research demo. A pipeline that needs a second is something a coding team can argue about with Health IT.
Temperature 0.0 is the result people will skip
They swept temperatures. Annotation quality and valid JSON-ish output were most stable at 0.0. Several models stayed near-perfect across the range. Two did not. mistral-large:123b dropped from a median F1 of 0.981 at temperature 0.0 to 0.789 at 1.0. Invalid responses went from 0% to 30%. nemotron-mini:4b dropped from 0.863 to 0.734, invalid from 0% to 8%.
If you only read marketing cards, mistral-large is the “strong” model and a 4B is the toy. At temperature 1.0 the strong model is the one inventing 30% garbage. The paper’s practical move is boring: lock 0.0 for screening and validation. Törnberg’s annotation advice already pointed at low temperature for deterministic labels. This study is the local-model version of that advice with a table.
Do not export “always 0.0” into a chatbot. This is a binary label on a clinical note, not a discharge summary in the patient’s voice. Sampling is a bug when the output is present/absent.
The F1 table is not a clean win
Validation F1, 750 notes each:
- vomiting 0.985 (95% CI 0.972–0.995)
- nausea 0.979 (0.966–0.989)
- dysuria 0.824 (0.734–0.896)
- diarrhea 0.814 (0.749–0.868)
Rule baseline: vomiting 0.913, nausea 0.724, dysuria 0.705, diarrhea 0.853.
Paired tests favored the LLMs for nausea and vomiting. Confidence intervals included zero for diarrhea and dysuria. Read that again. On diarrhea the rule list’s 0.853 sits above the LLM’s 0.814. The fancy local model is not automatically the annotator. For the two GI symptoms with messy language, a negation-aware dictionary still has a job.
Errors, per the paper, came from operational criteria (what counts as the symptom), temporal hedges, inconsistent documentation, missed mentions, and five mistakes in the reference labels. Five gold errors in a 1,100-note set is a reminder that “human baseline” is also a pipeline with bugs. They say so. Believe them, and do not treat 0.985 as a clinical device.
They also say the whole thing must be re-run per concept and per documentation culture. A German ED note is not your English discharge summary. gemma2:27b is not a general hospital brain. It won nausea in this screen.
The vulnerability paper is the same warning in code
Sabrina Kaniewski’s group, covered on September 30, reproduced six open-source retrieval-augmented vulnerability detectors. Most of the original papers leaned on proprietary models, private datasets, or private metrics. Under one shared dataset, one metric suite, and one open-weight pool, published wins often failed to transfer. Reproducibility varied. The underlying model mattered more than the RAG diagram.
That is the hospital paper’s cousin. If your method only works when the model is a closed API, you do not have a method you can hand to a hospital that is not allowed to send notes out. If your F1 only exists at temperature 0.7 with a private judge, you have a demo. The RAG study is not about ED notes. It is about the habit of reporting a number that dies in a rerun.
Hugging Face already retired the Open LLM Leaderboard in March 2025 after scoring more than 13,000 models, in part because static tests push people to climb the wrong hill. Composite indices and arenas filled the vacuum. A hospital still cannot pick a model off Arena Elo for nausea in German. It can steal this paper’s shape: small gold set, freeze the prompt, sweep temperature, screen locally, validate on a held-out slice, keep the rule baseline in the table.
What a research team should copy on Monday
Do not start with nine models. Start with the label book. If “diarrhea” in your notes includes “loose stool after bowel prep,” the model will look worse than a rule that ignores prep days, or better than a rule that cannot. The paper’s error list is a labeling meeting, not an ML meeting.
Then freeze the prompt. If you keep editing it while you shop models, you are fitting the prompt to the winner. They used the exploration set for that and then stopped. Copy the stop.
Sweep temperature on one symptom, not on the whole catalog. Their nausea sweep is enough to justify 0.0. If your stack only works at 0.8, you are not doing annotation. You are doing writing.
Keep a dumb baseline. The diarrhea row is the entire argument for that. An LLM screen that cannot beat a negation list is a GPU bill.
Report invalid-output rate next to F1. mistral-large’s 30% at temperature 1.0 would have been laundered out of a leaderboard that only shows the successful parses. Count the refusals and the broken JSON as failures. They are.
Publish the screening rule before you look at validation. Highest F1, then shorter inference as a tie-break, is what they wrote. Write yours. Otherwise the 27B that looks nice in a slide deck will win by default.
The 1,100-note set is German, single-center, emergency department. Nausea and vomiting are high-frequency, often explicit. Dysuria and diarrhea hide behind hedging, timing, and “after contrast” notes. That is why F1 splits. A model that looks brilliant on vomiting is answering the easy column. The paper’s screening chose different weights per symptom, which is the unglamorous conclusion: you may not get one hospital model. You may get four.
Positive predictive value is where diarrhea and dysuria sagged, even with high specificity. In English: the model was decent at saying “no,” worse at saying “yes” without collecting extra cases. A registry that pages clinicians on every positive cannot live on 0.81 F1. A registry that uses the model to sort the queue, then has a human confirm the yeses, can. Put the operating point in the protocol before you celebrate the table.
Bootstrap CIs at patient level are in the paper for a reason. Notes cluster. If you treat 750 rows as 750 independent people you will quote a tighter interval than you earned. Copy their interval style or do not quote a number in the slide.
The open-source implementation is the other deliverable. A methods paragraph without code is how RAG4SVD papers become unreproducible. If you fork this pipeline, keep the temperature sweep, the invalid-output counter, and the rule baseline in the same notebook. Deleting those three is how the next paper gets a pretty F1 and a dead rerun.
What this does not license
It does not license shipping an unattended coder into the chart. F1 on four binary symptoms is not a diagnostic model. It is a way to pre-label a registry so humans spend time on the disagreements. The five reference errors are the rate limit.
It does not license “open weights are safer.” Local deployment avoids a vendor cloud. It does not avoid a leaked box, a bad prompt, or a model that invents a symptom. The privacy win is real and separate from the F1 win.
It does not license dropping rules. On diarrhea, the rules were better. Hybrid is the adult version: LLM where it beat the list, list where it did not, human on the CI that includes zero.
And it does not license copying gemma2:27b into an English ICU. The paper’s last sentence is the one to tattoo on the README. Re-evaluate for each target concept and documentation setting. The open-source implementation is a workflow, not a weight file with a medical degree.
If you needed a research story this week that is not another 2.8-point average, this is it. A hospital locked temperature at zero, kept the rules in the table, and still lost a row. That loss is the part worth replicating.
One operational number still sitting in the methods: after the temperature lock, each of the nine models ran once on each symptom’s 250-note development set. No ten-seed average at that stage. The exploration set had the 10-iteration IQR work. Mixing those two protocols is how a later team “fails to reproduce” a number that was never a mean. Write which pass was a single shot.
The 100-note prompt slice is also easy to undersize. If your concept is rarer than nausea in an ED, 100 notes will not show you the hedge language. Steal the split ratios only if your prevalence looks like theirs. Otherwise grow the gold set until the development slice has enough positives to make F1 mean something. A 0.98 on 12 positives is a coin.