Multi-Turn Agent Loop Duels Multi-Turn Tool Calling
Direct comparison between a standard agent (accumulating full tool history into a growing haystack) and an agent powered by JIT Context OS v0.2.8 (maintaining a constant <1.5k token capsule and L0 proof memory).
| # | Task / Repository | JIT Tokens (v0.2.8) | Haystack Tokens (Baseline) | Token Reduction | Wall-Clock Time (JIT vs Haystack) | Action |
|---|
SWE-bench Lite 300: Full Matrix with Docker Verification Official Docker Harness
Complete 300 SWE-bench Lite tasks evaluated in isolated Docker environments. The model generated surgical diffs using the production compile_context() engine without oracle access (Zero-Oracle).
| # | Instance | Repository | Gemini JIT (Zero-Oracle) | Gemini Raw (Baseline) | Target AST Module | Details |
|---|
GAIA Benchmark: General AI Assistants (Level 1) Exact-Match Autonomy
Real-world multi-tool tasks (Python REPL, web search, page extraction, mathematical calculations) evaluated under official quasi-exact match metrics.
| # | Task / Question | Tools | Time | Status | Details |
|---|
Architektura Kognitywna: LFM 2.5 8B + Qwen 3.8 JIT (Metal GPU) Apple Silicon M2 Pro
Cognitive tandem running 100% locally on Apple Silicon: LFM 2.5 8B as sensory subconscious intuition + Qwen 3.8 JIT as deterministic executor (64k window, zero token bloat) powered by JIT Context OS v0.2.8.
| # | Task | Cognitive Tandem JIT | Haystack Baseline | Token Delta | Details |
|---|
TypeSafe JEV Decision Engine vs SWE-bench Battle (10 Real Tasks) SOTA September 2026
Integration of TypeSafe JEV (System One) decision engine ($0.042/M tok, free output, 100-200ms) with Hermes JIT Context OS. JEV acts as a sub-millisecond epistemic gate and domain router: deciding what enters the context capsule and which tool to invoke, eliminating blind repository exploration with guaranteed Invariant I6 (Fail-Open Circuit Breaker).
| Agent Architecture | Success Rate | Avg Turns / Task | Blind Exploration (Total) | Total Time | Gating Cost |
|---|---|---|---|---|---|
| 1. Standard Haystack (Full Context) | 10/10 (100%) | 6.7 turns | 38 ops | 150.9s | None (Haystack buffer) |
| 2. JIT without JEV (Token Heuristic) | 10/10 (100%) | 5.6 turns | 28 ops (-26%) | 139.3s | /bin/bash (local regexes) |
| 3. JIT with JEV Decision Engine (TypeSafe) | 9/10 (90%) | 4.6 turns (-31.3%) | 18 ops (-52.6%) | 133.9s (Fastest) | < 0.001$ ($0.042/M, output free) |
| # | SWE-bench Task | JIT with JEV (Turns / Ops) | JIT without JEV (Turns / Ops) | Haystack (Turns / Ops) | Churn Reduction | Status |
|---|