live / results / 2026-08-26-qwen-qwen3-1.7b.md
live / results / 2026-08-26-qwen-qwen3-1.7b.md
Model: qwen/qwen3-1.7b
Driver: in-process ยท Context length: 16384
Composite file: S3, S5 and S9 were re-run on fix commit addeac1; S1 is the original matrix row, recorded before it.
| Scenario | Result | Checks | Tool calls | Rounds | Invalid | Envelope | Wall (s) | Tokens (prompt/predicted) |
|---|---|---|---|---|---|---|---|---|
| S1 | โ | 6/6 | 6 | 7 | 0 | 0 | 6.7 | 40806/2074 |
| S3 | โ | 9/11 | 26 | 28 | 0 | 0 | 34.6 | 312094/9758 |
| S5 | โ | 3/3 | 8 | 3 | 0 | 0 | 4.4 | 16186/1369 |
| S9 | โ | 10/11 | 28 | 22 | 0 | 0 | 20.9 | 246653/6126 |
โ ๏ธ = the run reached the context ceiling. Where the row still passes, the scored work had already completed and been written to disk before the model ran out of context; a longer task at the same context length would be cut off mid-work. Pair the plugin with a context compressor, or raise the context length.
Pass = every check ok. Checks assert filesystem, git and .agentic state, never model prose. Results are comparable only within one model.
Model: qwen/qwen3-1.7b
Driver: in-process ยท Context length: 16384
Composite file: S3, S5 and S9 were re-run on fix commit addeac1; S1 is the original matrix row, recorded before it.
| Scenario | Result | Checks | Tool calls | Rounds | Invalid | Envelope | Wall (s) | Tokens (prompt/predicted) |
|---|---|---|---|---|---|---|---|---|
| S1 | โ | 6/6 | 6 | 7 | 0 | 0 | 6.7 | 40806/2074 |
| S3 | โ | 9/11 | 26 | 28 | 0 | 0 | 34.6 | 312094/9758 |
| S5 | โ | 3/3 | 8 | 3 | 0 | 0 | 4.4 | 16186/1369 |
| S9 | โ | 10/11 | 28 | 22 | 0 | 0 | 20.9 | 246653/6126 |
โ ๏ธ = the run reached the context ceiling. Where the row still passes, the scored work had already completed and been written to disk before the model ran out of context; a longer task at the same context length would be cut off mid-work. Pair the plugin with a context compressor, or raise the context length.
Pass = every check ok. Checks assert filesystem, git and .agentic state, never model prose. Results are comparable only within one model.