More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Pav’s follow-up to the “Needles and haystacks” series ran 26 combinations of Claude-4.6/4.7/4.8 and GPT-5.4/5.5 at four reasoning levels (low, med, high, xhigh) across two FreeBSD/OpenBSD vulnerability labs. Each model-effort pair was tested 20 times on both whole-file and function-level inputs—stacking up 2,080 runs per iteration and costing roughly $2,300 this round (about $9,200 total so far). He found that GPT-5.4 at medium or high effort outperformed everything else, scoring a normalized mean of 0.417 with 15 percent full solves and 76.2 percent partial hits on OpenBSD’s TCP sack bug, and even higher on the NFS vuln.
A four-LLM council voting system pushed accuracy to 86.2 percent unanimous decisions, dropping no-majority cases to just 2.8 percent. Odd-number councils likely beat evens. Most models flagged parts of the bug chain 70.8 percent of the time, but none cracked every link—full solves sat at 1.9 percent overall. GPT-5.5-med scored better than its high or xhigh runs, while low-effort modes trailed every time. Function-level prompts boosted performance dramatically compared to feeding the entire source file.
Content filtering spiked with higher reasoning settings, though Claude-4.7-1m had unusually low filter rates this iteration compared with past runs. Only the Claude variants cited CVEs by name. Pav suspects that smarter or “higher reasoning” modes sometimes overthink and miss simple patterns. All experiment code and raw JSON outputs live in the parsiya/mythos-bench-copilot repo for anyone wanting to replicate or drill deeper.
Questions about this article
No questions yet.