Live backend detected — the GPU worker is ready. Switch to LIVE
Helix × SmolLM2 MODEL PLAYGROUND SmolLM2-360M-Instruct · 360M · 2024 instruct chat model · GPT-2-XL switchable
PREVIEW · MOCK TELEMETRY (no GPU connected)
Models & paths Proof & attestation
Helix — the verifiable execution layer

Watch a modern instruction-tuned chat model think on a stack you can rebuild from 299 bytes.

SmolLM2-360M-Instruct is the headline model here — a 2024 Llama-architecture chat model that really follows instructions through its own ChatML template (recorded reply to “What is the capital of France?” → “The capital of France is Paris.”, verified token-for-token 8/8 by the oracle). The older GPT-2-XL base model is switchable too, and the offline preview replays it. Either way the live thing being proven is the verified compute underneath — every layer and kernel comes from the from-raw Helix toolchain.

SmolLM2-360M-Instruct · 360M · 32 layers · 2024 Llama arch 8 kovc-emitted kernels fp32 · greedy · GPT-2-XL (1.5B · 48 layers) switchable live pacing is intentionally slow; the pitch is trust, not speed
Conversation — text completion SmolLM2-360M-Instruct · fp32 · greedy
Give GPT-2 some text to continue
Pick a seed below or type your own. GPT-2-XL will continue it token-by-token while the 48 transformer layers and kovc kernels light up on the right. It is a base model — expect continuations, not answers.
token shade = p(chosen token) — real data from the live logits, “what the model actually considered”; click a token for alternatives <25% 25–45% 45–70% 70–90% ≥90%
Conversation = repeated completion with carried context. Each turn re-sends the conversation so far as one completion prompt — the model itself is stateless between requests: a 2019 base completion model, not an assistant. The live server caps the prompt at ~320 tokens (--max-ctx); when the carried text would blow that budget, the oldest text is cut first and the page says so.
Enter ↵ to run · Shift+Enter for a newline
GPU busy — one generation at a time (single-flight; the server keeps no queue, so the page just waits politely and retries).
Honest residuals: fp32-only · complete-to-PTX-not-SASS · single GPU (sm_86) · base-model-not-assistant · oracle-shares-spec · never-claimed-AGI. This is a demonstration of verifiable execution — not a claim of model quality, speed records, or full-GPU verification. No live parity verdict appears in this chat.