Multi-checkpoint voting ensembles of learned context-pruning models paradoxically underperform their best single-voter constituent. We formalize this: under k-of-N drop voting, the ensemble eviction indicator equals the k-th order statistic of the per-voter indicators — the ensemble collapses to its weakest member on every stratum. Three mechanisms resolve the paradox: asymmetric loss modulation during training (λ=3.0), regex-based inference overrides, and C3 self-distillation with a stronger teacher. We train 17 kompress models (149M-param ModernBERT) for $38.95 total, achieving a Pareto-optimal 0.955 heretic-exact at 15% compression with kompress-v8.
| Version | Heretic | Keep | Note |
|---|---|---|---|
| v2-base | 0.975 | 0.897 | precision ceiling |
| v8 ★ | 0.955 | 0.854 | production |
| v16 | 0.972 | 0.972 | endpoint (10x) |
| v17 | 0.963 | 0.963 | tradeoff (5x) |
GitHub · Models (18) · LoopKit · Eval Space · Blog
← back to kompress.vaked.dev