bench / results / 2026-08-24-meta-llama-3.2-3b-compact.md
bench / results / 2026-08-24-meta-llama-3.2-3b-compact.md
Model: meta/llama-3.2-3b
Variant: compact
Settings: {"limitTokens":500,"floorTokens":400,"keepRecentTokens":300,"chunkTokens":800,"summarizerMaxTokens":400,"consolidationBudgetTokens":600,"consolidationKeepNewest":2,"mergeMaxTokens":400,"maxExcerptChars":32000,"quizMaxTokens":120}
| Fixture | Recall (compacted) | Recall (consolidated) | Reduction | Chunk latency mean/p50/max (ms) | Structure | Cache |
|---|---|---|---|---|---|---|
| coding-agent | 5/5 | — | 61% | 564.7/555/648 | ✅ | ✅ |
| research | 1/3 | — | 42% | 832/832/832 | ✅ | ✅ |
| contradictory-updates | 1/3 | — | 37% | 276/276/276 | ✅ | ✅ |
| tool-spam | 3/3 | 3/3 | 68% | 466.8/380/753 | ✅ | ✅ |
Recall = seeded facts answered correctly from the compacted view alone. Results are comparable only within one model; see bench/README.md.
Model: meta/llama-3.2-3b
Variant: compact
Settings: {"limitTokens":500,"floorTokens":400,"keepRecentTokens":300,"chunkTokens":800,"summarizerMaxTokens":400,"consolidationBudgetTokens":600,"consolidationKeepNewest":2,"mergeMaxTokens":400,"maxExcerptChars":32000,"quizMaxTokens":120}
| Fixture | Recall (compacted) | Recall (consolidated) | Reduction | Chunk latency mean/p50/max (ms) | Structure | Cache |
|---|---|---|---|---|---|---|
| coding-agent | 5/5 | — | 61% | 564.7/555/648 | ✅ | ✅ |
| research | 1/3 | — | 42% | 832/832/832 | ✅ | ✅ |
| contradictory-updates | 1/3 | — | 37% | 276/276/276 | ✅ | ✅ |
| tool-spam | 3/3 | 3/3 | 68% | 466.8/380/753 | ✅ | ✅ |
Recall = seeded facts answered correctly from the compacted view alone. Results are comparable only within one model; see bench/README.md.