Fosniedocsv0.7

Automatic model routing

Configure local, cloud and mixed model pools, Auto strategies, judges, effort, limits and failure recovery in Fosnie Core and Enterprise.

Fosnie v0.7.0 can select a model for a request from several configured local or cloud providers. Choose a model directly, keep Default, or select Auto. Auto combines the task's requirements with current model profiles, permissions, routing rules and the run's remaining allowance.

This is an open-source Core capability, also available in Enterprise. A local routing model is optional: a fully cloud-based installation can use a cloud judge. Enterprise restrictions can narrow the permitted models and source access; they cannot widen a mandatory Core restriction.

Auto chooses from the capabilities and quality information configured for your installation. It does not know that a model is better merely because its name is newer. Operator declarations, successful connection tests and independent quality evidence are different things. Automatic routing is shipped; a guarantee of the best answer or financial savings is not.

Selection, strategy and rollout

These are separate settings. Selecting Auto in the composer does not mean that every request runs through a judge or uses the most expensive model.

SettingValuesEffect
User selectionDefault, Manual, AutoKeep the configured default, pin a visible provider, or let a policy select the executor.
Auto strategyRules, Classifier, HybridUse deterministic rules, a bounded judge call, or rules followed by classification when needed.
ObjectiveEconomy, Balanced, QualityChange preferences inside the eligible pool. An objective cannot override permissions or capabilities.
Policy rolloutOff, Shadow, ActiveUse the checked default, compare a proposed choice while the default answers, or execute the proposal.
ConstraintsPool, locality, capabilities, budgetsHard filters shared by the applicable inference purposes and any permitted alternate.

Manual bypasses the routing judge and retains the selected provider. It still obeys access, locality, capabilities, context and budget limits. Default uses the effective configured default for an adopted routed caller. Neither silently changes into Auto.

An omitted selection inherits the saved conversation choice. Explicit Default resets Manual/Auto. Local-only and effort are saved independently. A saved unavailable Manual model produces an error in the routed path; it does not secretly use another provider or an environment key.

How a decision is made

  1. Authenticate the original actor and check the caller, agent, project and permitted source scope.
  2. Resolve the saved selection and policy, and intersect deployment, task and Enterprise restrictions.
  3. Build bounded requirements from the request and permitted conversation context.
  4. Apply rules and, when the selected strategy requires it, call the assigned judge once.
  5. Filter current profiles for locality, tools/JSON/vision, context, output and effort compatibility.
  6. Rank eligible model/effort pairs using the objective and available profile information.
  7. Check the assembled execution context and reserve the shared allowance before dispatch.
  8. Record the actual executor and accounted usage; keep the model pinned through its tool loop.

The judge classifies the task. It does not choose an executor by guessing which model name is strongest, execute tools, or route itself through Auto. Model selection stays in the server's policy engine.

The request's length alone is not its difficulty. A short proof, a long extraction and a continuation of a complex conversation have different requirements. Unknown tasks, low confidence, truncated classification input or judge failure use a compatible checked default within the same remaining allowance. If no compatible default exists, the request fails explicitly.

Configure an installation

1. Connect providers

Add the intended models through Providers. Each LLM row represents an endpoint, model and credential together. Use the adapter appropriate to that endpoint:

Profile adapterEndpoint shape
openai_chatOpenAI-compatible Chat Completions, including compatible Ollama, llama.cpp or vLLM endpoints.
openai_responsesNative OpenAI Responses.
anthropicNative Anthropic Messages.
geminiNative Gemini GenerateContent.

Native cloud adapters do not require you to put an OpenAI-compatible gateway in front of them. Do not select a native adapter for an incompatible proxy URL. An explicitly empty key sends no Bearer header; it does not inherit another provider's key. Routed inference does not follow HTTP redirects with credentials.

Providers retain their existing deployment/personal ownership and BYOK rules. Adding routing does not make another user's key visible or permit personal users to edit a shared provider's rating.

2. Save current model profiles

Open Providers → Model routing and save a profile for each intended executor and judge. Declare:

  • adapter and the exact configured model;
  • context/output bounds and supported effort values;
  • tools, structured output and vision capabilities where known;
  • locality and whether a deployment administrator has attested it;
  • overall/task tiers, exclusions and known latency/tariffs;
  • whether the model may participate in Auto and its optional Auto effort preset.

Unknown support remains unknown. A familiar name does not prove tool support, schema support, context size or quality. Use conservative bounds appropriate to the deployed endpoint, not the largest advertised bound if that endpoint cannot serve it.

Example profile fields below are illustrative. Replace the UUID/model with your saved provider and only declare capabilities you have established for that deployment:

{
  "provider_id": "00000000-0000-0000-0000-000000000101",
  "provider_revision": 1,
  "revision": 1,
  "adapter": "openai_chat",
  "model": "your-local-model",
  "locality": "local",
  "locality_trusted": true,
  "quality_status": "declared",
  "tier": "balanced",
  "auto_enabled": true,
  "auto_effort": "none",
  "capabilities": {
    "context_tokens": 8192,
    "max_output_tokens": 2048,
    "tools": "declared",
    "structured_output": "unknown",
    "efforts": ["none"]
  }
}

quality_status: declared is an operator statement, not measured quality. A BYOK user cannot self-attest trusted locality for a shared deployment. Provider/model/endpoint changes invalidate old references and evidence; reload current revisions instead of overwriting a stale edit.

3. Create a policy

Choose a compatible checked default, strategy, objective, model pool and limits. For Classifier/Hybrid, explicitly assign the judge and its input/output/deadline bounds. There is no mandatory local judge or hidden fallback to an unspecified cloud judge.

This illustrative Hybrid policy starts in Shadow. All UUIDs must refer to profiles the saving actor may use:

{
  "id": "00000000-0000-0000-0000-000000000201",
  "revision": 1,
  "strategy": "hybrid",
  "rollout": "shadow",
  "objective": "balanced",
  "default_provider_id": "00000000-0000-0000-0000-000000000101",
  "classifier_provider_id": "00000000-0000-0000-0000-000000000102",
  "rules": [],
  "technical_fallback": true,
  "constraints": {
    "allowed_providers": [
      "00000000-0000-0000-0000-000000000101",
      "00000000-0000-0000-0000-000000000102",
      "00000000-0000-0000-0000-000000000103"
    ],
    "max_calls": 6,
    "max_output_tokens": 1024,
    "max_total_tokens": 32768,
    "max_duration_ms": 90000
  },
  "classifier": {
    "native_schema": false,
    "min_confidence": 0.75,
    "max_input_tokens": 4096,
    "max_output_tokens": 512,
    "local_timeout_ms": 30000,
    "cloud_timeout_ms": 30000,
    "answer_reserve_ms": 15000
  }
}

Do not copy sample limits without checking your workloads. A judge's confidence is not a calibrated probability. A Shadow judge still consumes allowance even though the default answers.

4. Preview, test and enable the caller

The no-inference preview shows declared requirements and exclusions. Connection probes and judge tests make real calls when launched and do not establish answer accuracy.

After inspecting the configuration, enable the intended caller and set the policy to Active. Caller gates are independent: enabling browser chat does not also enable API, voice, research or workflow routing. Upgrades start new gates disabled and preserve saved intent; they do not turn every old conversation into Auto.

Rules, Classifier and Hybrid

Rules makes no classifier call. A matching terminal selection chooses its permitted target; no match uses the compatible configured default. It does not pick a random cheap model.

Classifier makes one bounded call to the assigned judge after hard checks. A malformed response, timeout, cancellation or insufficient allowance never creates an unbounded classification loop. Cancellation stops; other eligible judge failures use the checked default when possible.

Hybrid lets a clear terminal rule finish the selection. Otherwise it classifies. Conservative built-in features do not treat the word “weather”, “code” or the message's length as sufficient proof of task difficulty when context conflicts.

Rule conditions can use task, purpose, language, token bounds, tools, vision, freshness, structured output and uncertainty, with bounded nested ALL/ANY conditions. Actions can select a provider/effort, require a minimum tier, narrow the main executor pool or request classification.

A rule's executor pool applies to the main executor and its eligible technical alternate. The mandatory policy pool applies to every inference purpose, including judges, helpers and fixed services. A rule does not create permission to search the web, read documents or call a forbidden model.

Local, cloud and mixed pools

InstallationTypical configuration
Fully localTrusted local executor, local judge if needed, local helpers/fixed services and mandatory local_only. Rules can avoid a judge entirely.
Fully cloudCloud executors and a permitted cloud judge. No local LLM runtime is required.
MixedLocal and cloud executors, with an explicitly chosen local/cloud judge. Use objectives, rules, capabilities and evidence to control selection.

Mandatory Local inference only applies to inference across the run, including judge/helpers/fixed services and any alternate. A cloud judge cannot receive a request simply to decide which local executor to use under that restriction.

Local inference and source/network access are separate. A local model may use authorised web search; local-only does not grant that tool or make search-engine queries private to the host. Conversely, restricting the main executor to a local rule pool does not necessarily restrict explicitly assigned helpers: use mandatory locality if all inference must stay local.

Reasoning effort and auxiliary stages

Effort values are model/adapter-specific. Declare the values the deployed endpoint supports; Fosnie does not silently remove an unsupported effort.

For Auto execution, precedence is:

  1. explicit request effort;
  2. explicit terminal-rule effort;
  3. the model's auto_effort preset;
  4. the provider default.

Conflicting explicit request/rule efforts fail. Manual and Default do not inherit Auto presets; they use provider defaults unless the request supplies an effort. Auto Off/Shadow checked-default execution uses the selected default's Auto preset.

The primary, tool loop and final answer retain their assigned model and effort. Explicit utility assignments can use another permitted model for tasks such as titles/summaries. Each helper has its own purpose/capability checks and shares the original allowance. The classifier's effort is separately configured; it does not inherit the executor's preset.

Embeddings, rerank, OCR, STT, TTS and verification use fixed role profiles. Routing does not silently replace an embedding model/index or migrate a pinned fixed service during a run.

Web search and lightweight Quick mode

A routed web agent needs the existing web tool permission, egress permission and compatible executor. Search does not bypass source restrictions. Standard search retains its web plan/grade and fixed-service dependencies.

Administrators can enable Web search: use snippets for Quick searches (web_search.quick_snippets_only). The default is false. Effective Quick requests then use bounded search-engine excerpts, preserving URLs, dates, domain rules and citations, without fetching pages, OCR, reranking inference or a web planning model.

Cap the agent's maximum web depth at Quick if those unused fixed services are absent. An agent that permits Standard/Deep keeps the existing full dependencies. Changing the configuration cannot enter an unplanned fixed service.

The digest explicitly marks snippet-only, partial coverage. A forecast is not a current observation, and an excerpt does not prove exhaustive coverage. Empty results say that the fact could not be established. This mode reduces the work performed; it is not a factual-accuracy guarantee.

Failures and recovery

ConditionBehaviour
Judge unavailable, invalid JSON or low confidenceUse the compatible checked default if allowed and sufficient allowance remains.
Connection failure, timeout, transient provider failure or confirmed quota/payment errorOne technical alternate may be attempted at an eligible pre-output safe point when enabled.
Rejected credentials/access or invalid requestSurface a diagnostic; these do not authorise the same automatic retry path.
Shared ML service unavailableStop with an infrastructure diagnostic; changing providers does not repair that service.
No eligible model, stale configuration or exhausted allowanceStop explicitly; do not relax restrictions or use a hidden environment key.
CancellationStop and do not retry.
Output/tool/action already emittedRetain the partial response and stop; do not splice another model's answer or replay actions.

Enable Technical fallback before output on an Active Auto policy. The stateless text API and eligible plain browser/paired-desktop remote text turn can make at most one alternate attempt before tokens, reasoning, tool calls or other model events. The same question, actor, turn, cancellation state, deadline and spent allowance remain in force; no second judge or duplicate user message is created.

The alternate is rechecked against current policy/profile revisions, visibility, locality, capabilities, rule pool, effort and remaining allowance before credentials are resolved. Tool/RAG/fixed/artefact/voice/durable paths do not gain automatic replay. Manual/Default do not silently switch models.

Provider failures cause revision/purpose-bound cooldowns: normally 30 seconds for transient failures and 300 seconds for confirmed quota/payment/credential failures, or a bounded positive Retry-After. Malformed completed judge replies cool the judge purpose. Credential equality is used only where known; different keys/accounts or replicas are not magically identified as one billing account.

The UI uses safe typed descriptions, preserves partial content and failures after reload, and shows the actual alternate where visible. Prepare retry restores the original question to the composer; it does not send it. Old clients still receive the established chat error frame.

Budgets and actual usage

Judge, executor, helpers and retry share one run allowance. Calls reserve before dispatch. Failed, cancelled or unconfirmed work is not assumed free. A valid complete report can release unused reservations; unknown usage retains a conservative bound or causes refusal where a required bound cannot be proved.

Known tariffs support estimates/caps, not invoices. Unknown prices and spend remain unknown. Fixed services have separate call/byte/deadline accounting; hard monetary/token caps without proven bounds reject those dispatches. Objective Economy is a preference among eligible candidates, not a promise of lower cost on every request.

Actual model/effort is recorded only after accounted dispatch. Shadow proposals are distinct from the answering model. History hides model details that the caller may no longer view; provider keys and native error bodies are not product diagnostics.

API integration

For OpenAI-compatible clients, configure the API Auto alias in Providers → Model routing. The initially disabled alias is fosnie/auto; choose another name if a provider already uses it. Existing provider names retain their meaning, and agent/ is reserved.

The alias, the independent public-routing gate and the existing public API entitlement must all permit the call. Activation does not change policy rollout. Discover the configured name through GET /v1/models and routing capabilities through GET /v1/routing.

{
  "model": "fosnie/auto",
  "messages": [{"role": "user", "content": "Explain this algorithm."}],
  "max_completion_tokens": 1024,
  "stream": false
}

Send this to POST /v1/chat/completions with a Fosnie platform API key, not the provider's upstream key. Raw API routing is stateless text; it does not add implicit tools or RAG. An explicitly adopted agent API has a separate bounded tool/source scope.

The browser/paired-desktop wire intent is capability-discovered chat.routed.send:

{
  "version": 1,
  "type": "chat.routed.send",
  "request": {
    "chat_id": null,
    "agent_id": null,
    "project_id": null,
    "content": "Explain this algorithm.",
    "model_selection": {"mode": "auto", "objective": "balanced"},
    "options": {"local_only": false, "effort": null}
  }
}

This is intent, not a client-supplied execution plan. Management routes are session-authenticated and enforce ownership/effective provider-management permissions; stale revision writes are conflicts rather than silent overwrites.

Supported scope and boundaries

Entry pointCurrent scope
Browser / paired desktop remote chatGeneral remote conversations, saved intent/options and actual metadata. Existing agent/source permissions apply. Attachments, local workspace/device tools and branch continuation are not adopted by this entry.
Raw API / virtual Auto aliasStateless text JSON/SSE, without implicit tools/RAG.
Adopted agent APIExplicit supported agent/library/web scope; unsupported tools or approval/resume fail before actions.
Document, research and workflow jobsIndependent opt-in, original actor/source access, pinned selection, fenced leases and shared accounting. Existing delivery permissions still apply.
Voice / telephoneIndependent opt-in, fixed STT/TTS and bounded batch audio; unsupported realtime/speculation/actions remain excluded.
Completed-answer exportSupported bounded formats without another LLM call; native PDF rendering has a separate prerequisite.

Unsupported combinations fail explicitly. They do not drop tools/effort, truncate context silently, switch to legacy execution or repeat stateful side effects.

Evidence and optional quality escalation

Optional learned_ranking and quality_escalation start disabled. Reviewed ranking can use comparable held-out records tied to the current provider/profile and call purpose. Admission requires at least 20 independent dialogue families per comparable candidate and the existing quality-floor checks. Synthetic, stale, unreviewed or incomparable results do not affect learned ranking.

Quality escalation is a bounded structural-answer check for eligible stateless, non-streaming Auto API work before results are exposed. It shares the same retry allowance with technical fallback. It is not a general semantic judge, an ensemble, online exploration or model switching inside a tool loop.

The release's operational tests establish routing, accounting, permissions and failure behaviour. They do not admit independent model-quality evidence or prove that the chosen model is always best. Evaluate your actual workloads separately before making quality or savings claims.

Upgrade, disable and troubleshoot

Use matching backend, ML and frontend components and the normal SQLx migration runner. v0.7.0 includes additive routing migrations through 0138. Back up before upgrading, review profiles/policies and enable only intended caller gates. See Upgrades & TLS and Backups & DR.

To suspend automatic selection, save policy Off with a compatible checked default. Hard limits remain; settings/history stay saved. To stop a caller entirely, disable its independent gate. Adopted conversations cannot silently downgrade to legacy. Explicit re-enable applies current permissions/revisions to new plans.

Running an older binary directly against the new schema is not certified. Restore a verified pre-upgrade backup into an isolated installation before switching back, using application artefacts matching that snapshot. New P12 fields may be rejected by older strict readers; removing a few columns is not a safe downgrade procedure.

SymptomCheck
Auto unavailableCaller capability/gate, supported conversation provenance and effective permissions.
No eligible modelCurrent profiles, locality/allowlist, required tools/JSON/context/effort and remaining allowance.
Judge skipped/default answeredCompatible candidate count, confidence, context/deadline reserves, judge locality/access and failure reasons.
No fallbackActive Auto, enabled technical fallback, pre-output plain/stateless safe point, permitted alternate and sufficient allowance.
Effort rejectedExact adapter/model support and conflicts between explicit request and rule.
Quick web needs OCR/rerankEnable snippet mode and cap agent depth at Quick; otherwise full dependencies remain.
Configuration conflictReload current revisions after provider/profile/policy edits.
Alias unavailableExisting-name collision, alias activation, public-routing gate and API entitlement.

The v0.7.0 deployment templates also adopt Fosnie names. Legacy PAI__*/PAI_CONFIG_FILE remain fallback aliases, with FOSNIE__* winning when both are supplied. Systemd/launchd service names and filesystem paths require host changes or explicit path overrides; the PostgreSQL role/database is not automatically renamed. The release notes list the migration details.

Was this page helpful?

On this page