bench / results / 2026-08-24-openai-gpt-oss-20b.md
bench / results / 2026-08-24-openai-gpt-oss-20b.md
Model: openai/gpt-oss-20b
Settings: {"limitTokens":500,"floorTokens":400,"keepRecentTokens":300,"chunkTokens":800,"summarizerMaxTokens":400,"consolidationBudgetTokens":600,"consolidationKeepNewest":2,"mergeMaxTokens":400,"maxExcerptChars":32000,"quizMaxTokens":120}
| Fixture | Recall (compacted) | Recall (consolidated) | Reduction | Chunk latency mean/p50/max (ms) | Structure | Cache |
|---|---|---|---|---|---|---|
| coding-agent | 5/5 | — | 33% | 1739.7/1617/2045 | ✅ | ✅ |
| research | 2/3 | — | 34% | 1453/1453/1453 | ✅ | ✅ |
| contradictory-updates | 3/3 | — | 10% | 1251/1251/1251 | ✅ | ✅ |
| tool-spam | 1/3 | 1/3 | 61% | 1131.4/1132/1639 | ✅ | ✅ |
Recall = seeded facts answered correctly from the compacted view alone. Results are comparable only within one model; see bench/README.md.
Model: openai/gpt-oss-20b
Settings: {"limitTokens":500,"floorTokens":400,"keepRecentTokens":300,"chunkTokens":800,"summarizerMaxTokens":400,"consolidationBudgetTokens":600,"consolidationKeepNewest":2,"mergeMaxTokens":400,"maxExcerptChars":32000,"quizMaxTokens":120}
| Fixture | Recall (compacted) | Recall (consolidated) | Reduction | Chunk latency mean/p50/max (ms) | Structure | Cache |
|---|---|---|---|---|---|---|
| coding-agent | 5/5 | — | 33% | 1739.7/1617/2045 | ✅ | ✅ |
| research | 2/3 | — | 34% | 1453/1453/1453 | ✅ | ✅ |
| contradictory-updates | 3/3 | — | 10% | 1251/1251/1251 | ✅ | ✅ |
| tool-spam | 1/3 | 1/3 | 61% | 1131.4/1132/1639 | ✅ | ✅ |
Recall = seeded facts answered correctly from the compacted view alone. Results are comparable only within one model; see bench/README.md.