Planner
Router
Workers
Verify
Consensus
Synthesize
Multi-model orchestration

One prompt.Every model, coordinated.

Ask once — KernelFold routes your request across 15 frontier models from labs like Anthropic, OpenAI, DeepSeek and Moonshot, runs specialists in parallel, verifies the work, and streams back one answer. It learns per task and routes around outages automatically.

Free to start · no credit card · OpenAI-compatible API

Orchestrating models from

AnthropicOpenAIDeepSeekMoonshot AIZ.aiQwenMiMoAnthropicOpenAIDeepSeekMoonshot AIZ.aiQwenMiMo
The fleet

Fifteen models. Seven labs.One answer.

You never pick a model — the router does, per task, using live capability, cost, latency, and learned accuracy. These are the flagships it leads with — every plan gets all of them.

AnthropicFrontier

Claude Fable 5

Anthropic's Mythos-class flagship — the sharpest reasoner in the fleet

200K context

OpenAIFrontier

GPT-5.6 Sol

OpenAI's frontier lane — elite reasoning & agentic coding

400K context

AnthropicFrontier

Claude Opus 5

Deep reasoning & verification at million-token context

1M context

AnthropicStrong

Claude Sonnet 5

The Claude 5 workhorse — frontier-family quality, everyday speed

200K context

DeepSeekStrong

DeepSeek V4 Pro

Deep technical reasoning at great value

128K context

Moonshot AIStrong

Kimi K3

The fleet's top dedicated coder

256K context

Z.aiStrong

GLM 5.2

General work at million-token context, great value

1M context

OpenAIStrong

GPT-5.6 Terra

Frontier reasoning with native tool use & vision

400K context

See all 15 models

Availability is monitored continuously — if a model degrades mid-answer, KernelFold switches and keeps streaming. See live status

Why KernelFold

Built like an ensemble,feels like one model.

learn

A router that learns

Model choice blends capability, cost, latency, and historical accuracy — read live from the database on every request. The more you use it, the smarter it routes.

speed

Two speeds

KernelFold streams the single best model for latency. KernelFold Ultra runs the full ensemble — plan, route, parallel workers, verify, reflect, synthesize.

resilience

Resilient by design

Mid-stream drops and gateway 502s are retried with automatic model fallback. The user still gets an answer.

observe

Fully observable

Every pipeline stage emits an event. Watch the pipeline animate live; audit cost, latency, and routing rationale after.

verify

Independent verification

A separate model verifies each answer before consensus — confidence you can see, not vibes.

The method

Five stages.One answer.

  1. 01

    Plan

    The planner decomposes your request into tasks with dependencies.

  2. 02

    Route

    Each task is scored against every model on capability, cost, latency, and learned accuracy.

  3. 03

    Work

    Specialist workers run in parallel, each on the model best suited to its task.

  4. 04

    Verify

    An independent model checks the work — errors are caught before you see them.

  5. 05

    Synthesize

    Consensus resolves disagreements; one answer streams back, live.

Pricing

Simple plans.Tokens that map to real work.

Your plan is a monthly pool of tokens shared across every model and the API — a quick chat spends a few thousand, a deep multi-agent run more. Paid plans also have a rolling five-hour window and a fixed weekly quota: these add no tokens, they just keep one runaway agent loop from draining a month in an afternoon. Allowances reset monthly. When you run out, upgrade or wait for the reset; we never charge overages.

Free

Try the whole fleet, every month.

$0/month

1.5M tokens every month

Every month
1.5MYour allowance
Pace limits
none

The 1.5M-token monthly allowance is the only cap — no rolling windows, no weekly quota. When it runs out, upgrade or wait for the monthly reset.

  • Full access to all models & both tiers
  • ~150 quick chats or ~8 deep runs / month
  • OpenAI-compatible API & dashboard
Start for free
Most popular

Pro

For daily individual use.

$19/month

120M tokens every month

Every month
120MYour allowance
Per 5 hours
4MRolling
Per week
30MResets weekly

Up to 4M tokens in any rolling five-hour window and 30M per week, subject to your 120M monthly allowance.

  • Everything in Free, ~80× the tokens
  • Deep multi-agent runs as your default
  • Scoped API keys for your tools & editors
Get Pro

Max

For heavy and agentic workloads.

$49/month

360M tokens every month

Every month
360MYour allowance
Per 5 hours
7MRolling
Per week
90MResets weekly

Up to 7M tokens in any rolling five-hour window and 90M per week, subject to your 360M monthly allowance.

  • 360M tokens vs Pro's 120M — 3× as many
  • More headroom for Claude Code, Codex, and agent loops
  • Same model fleet, same API — just more of it
Get Max

Every plan uses the same models and the same API — plans differ only in monthly tokens and how fast you may spend them. Upgrade, downgrade, or cancel any time from your billing page.

FAQ

Frequently askedquestions.

Everything else lives in the docs — or ask us directly.

What is KernelFold AI?
KernelFold AI is a multi-model AI assistant. Instead of sending your question to a single model, it plans the request, routes it across multiple frontier models (such as Claude, GPT, Kimi, and DeepSeek), runs specialists in parallel, independently verifies their work, and synthesizes one best answer — streaming every step of the pipeline so you can see how the answer was produced.
Which AI models does KernelFold use?
KernelFold draws on a fleet of frontier models from multiple labs, including Anthropic's Claude (up to Claude Fable 5 and Opus 5), OpenAI's GPT (up to GPT-5.6), DeepSeek, Moonshot's Kimi, Z.ai's GLM, Qwen, and more. The router picks the best model for each task automatically and routes around any provider outage. The full catalog is on the Models page at kernelfold.com/models.
How is KernelFold different from ChatGPT or using one model directly?
A single model gives you one model's answer. KernelFold coordinates several: it decomposes hard questions into subtasks, sends each to the model best suited for it, has an independent step check the results, and merges them into a single verified response. For difficult or high-stakes questions this reduces individual-model blind spots and hallucinations.
Does KernelFold have an API?
Yes. KernelFold exposes an OpenAI-compatible API, so you can point existing OpenAI SDKs and tools at it by changing the base URL and key. You can call the full orchestration pipeline or route directly to a specific model. See the documentation for endpoints and examples.
Is KernelFold free to use?
You can start for free with no credit card required. Paid plans add higher usage limits and the full multi-model Ultra pipeline. Current tiers and limits are listed on the Pricing section of the homepage.
When I ask which model I'm talking to, what does KernelFold say?
KernelFold tells you exactly which model or models produced your answer. On the single-model tier it names the specific model that replied; on the Ultra tier it names every model that contributed to the synthesized answer. It never claims to be another company's assistant.
Start in under a minute

Stop choosing a model.Orchestrate all fifteen.

  • Chat in the browser or call the OpenAI-compatible API
  • Scoped API keys, usage dashboard, live pipeline view
  • Automatic failover — an outage never blanks your answer

Free to start · no credit card required

Works with any OpenAI SDK
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://www.kernelfold.com/api/v1",
  apiKey: process.env.KERNELFOLD_API_KEY,
});

const res = await client.chat.completions.create({
  model: "kernelfold-ultra",   // the whole fleet
  messages: [{ role: "user", content: "..." }],
  stream: true,
});