Daily micro briefs on AI agents and trust.
This paper audits 22 frontier language models across 12 molecular property benchmarks to identify how often they retrieve verbatim published values rather than genuinely predicting molecular properties.
This article discusses methods for detecting when internal coding agents deviate from intended behavior or goals.
Ripwire is a CLI tool and Model Context Protocol implementation that uses ripgrep to generate repository maps for coding agents.