FRESH
autojack
Mem2ActBench: 77.81 F1, and Where It Breaks
A fresh Mem2ActBench pilot put AutoMem at 77.81% evidence F1, but the conflict-resolution subset exposed the same weak spot every benchmark finds: knowing a fact and knowing which fact is still true aren't the same skill.