THEONES RESEARCH / BENCHMARK HUB
EMPIRICAL RUNTIME
SOTA Empirical Benchmarks • Zero-Oracle Invariant v0.2.8 • Mac Mini M2 Pro + Borg Cluster

Hermes JIT Context OS: Benchmark Hub

Catalog of physically measured performance for JIT Context OS (L0 SQLite WAL, L1 AST Working Set, L2 Distillation Cascade) and TypeSafe JEV System 1 Epistemic Gate. All benchmarks derived from authentic code executions without synthetic estimations.

Token Reduction Agent Loop
-70.4%

148k tok (JIT) vs 501k tok (Haystack) across 10 SWE-bench tasks.

Wall-Time Speedup Speedup
1.63x – 2.6x

315s with JIT vs 515s standard agent by eliminating prefill latency.

Local Apple Silicon Metal GPU
30.2 tok/s

Qwen 2.5 Coder 7B on Metal M2 Pro: -63.6% lower VRAM footprint.

Docker Verified Passes Official Harness
21 / 63

33.3% resolved tasks on applied patches in SWE-bench Lite.

Multi-Turn Agent Loop Duels Multi-Turn Tool Calling

Direct comparison between a standard agent (accumulating full tool history into a growing haystack) and an agent powered by JIT Context OS v0.2.8 (maintaining a constant <1.5k token capsule and L0 proof memory).

# Task / Repository JIT Tokens (v0.2.8) Haystack Tokens (Baseline) Token Reduction Wall-Clock Time (JIT vs Haystack) Action