I set out to defend a number, and found two of my own bugs instead

I set out to defend a number, and found two of my own bugs instead


New
Agentic AILocal LLMsDeveloper ToolsOpen Source

The pushback was fair

After I published the benchmark writeup, the comments that actually stuck with me weren’t the dismissive ones. They were the specific ones. “Sounds like a chat template issue.” “Why test old models - is that just picking a target that guarantees a bad number?” Both are the right instinct for anyone reading a benchmark that happens to favor the person who wrote it. I didn’t want to answer either one with a shrug.

So I checked. Properly.

Ruling things out, not just asserting them

The template question got a real answer: I pulled the exact instruction Ollama sends the model and diffed it, line for line, against the model creator’s own officially published chat template. Same wording, same tokens - Ollama’s version was actually stricter. Not a template bug. Then I went further than the question asked and installed vLLM from scratch on a GPU architecture it had never been tested on here, just to rule out “different serving software” too. Same failure, on completely different plumbing.

The “why old models” question deserved more than a defensive answer too. So instead of arguing, I ran the same six tasks all the way up the scale - 7B, 14B, 32B - and then jumped to a full model generation newer, and then to the actual most current release available. Not to cherry-pick a result. To find out.

The result I almost got to feel good about

And then something genuinely surprising happened: on the newest model, three separate competing tools all hit a perfect score. First time that’s happened anywhere in this whole series. If I’d stopped there, the honest story would have been “the gap is closing, and that’s worth saying plainly.”

Except my own number came back worse than everyone else’s. On a model that, by every other measure, was working fine.

Catching yourself is the actual work

That mismatch is the part I’m most glad I didn’t wave away. The easy move would have been to assume the model just didn’t like my approach as much - a plausible-sounding story, and wrong. Digging into the actual transcripts, it wasn’t the model at all. It was me. A tag-matching bug in my own parser that had nothing to do with reliability. Then, once I fixed that, a test-harness timeout that was punishing a bigger model for taking longer to think, not for being unreliable.

Two real bugs, both mine, both fixable, both hiding behind a number that looked like a finding.

Fixed both. Reran. Landed exactly where the field did - tied, not ahead, not behind.

What I actually took from this

The lesson isn’t “always be more rigorous,” which is true but useless as advice. It’s narrower than that: the moment a result surprises you in your own favor is exactly the moment to distrust it hardest, not publish it fastest. A number that makes you look bad and turns out to be a real bug in your own code is a much better outcome than a number that makes you look good and turns out to be a mistake you never caught.

Reliability - in the models, and in my own testing of them - turned out to be genuinely uneven in ways that weren’t obvious until I actually measured them instead of assuming either direction. That’s a less satisfying story than “everything’s broken” or “everything’s fine.” It’s also the true one.

Full technical writeups, if you want the receipts: the chat template investigation and the current-model result and the bugs it surfaced.

© 2026 Giuseppe Sirigu
Hi! I'm Giuseppe's AI assistant — ask me anything about his background 👋