Intelligence Brief
The first controlled eval of agent skills, Naming a test technique buys the framework, not th, The skill that did work was written quickly by som
3 of 4
packaged agent skills that underperformed giving the agent no instructions at all, in the first controlled measurement of skills — 26 conditions, 80 runs each
Today's Insights
The first controlled eval of agent skills — 26 prompt conditions, 4 skills, 80 runs each, implementing Zstd in Rust — found that 'Default (no additional instructions) does well above average' and that three of the four tested skills underperformed it, including the one from a repo with 250,000 stars.
Naming a test technique buys the framework, not the technique: agents asked to use formal methods wrote vacuous A-implies-A proofs in Verus, agents asked for property-based testing hammered invalid inputs against trivial properties, and one test of a four-bitstream feature submitted four identical bitstreams.
The skill that did work was written quickly by someone who says he has 'no feel for how to write a good skill' — the stated difference is that it nudges agents away from their default behaviour, while the underperformers read like tutorials. One was over 20,000 tokens.
Ken Thompson's trusting-trust attack is not a compiler problem: researchers built a complete one around GNU strip, a utility that never touches source code, and a single tampered strip in the NixOS binary seed backdoored almost every binary of a full graphical installer — surviving after the seed left the dependency closure.
Terrastruct is shutting down and open-sourced TALA, its diagram layout algorithm, under MPL-2.0. Its standout feature exists because of a measured model weakness: you can pin node coordinates and let the engine route the edges, since 'models can draw in 2D space well, but still struggle with routing'.
Video production now ships as twenty agent skills behind a router: HyperFrames turns HTML, CSS and seekable animations into deterministic MP4, installs with one npx command, and its own README warns that the skills registry blob can lag the repo by hours.
Context Mode moves context management to the MCP protocol layer — a subprocess sandbox where raw data never enters the conversation, 315 KB becoming 5.4 KB. Its most useful number is the enforcement gap: about 98% saved via hooks versus about 60% via instruction files, because instructions guide but cannot block.
Breeze TTS 2 tops the open-weight text-to-speech leaderboard at 1,215 Elo with sub-40ms time-to-first-audio and voice design from a plain-language description — but the weights are research-and-non-commercial-only and it needs a 12 GB CUDA GPU, so quality went up while deployability went backwards.
Trending Repos
No trending snapshot for this episode (GitHub trending JSON was not in the research bundle). Open live trending on GitHub, or check topic summaries in Today's Insights for named repos.
github.com/trending