Refonte Learning: Function-Calling vs Structured Outputs for Agent Design in 2026

Function-Calling vs Structured Outputs for Agent Design in 2026

Sat, Jun 27, 2026

Executive summary

Function-calling and structured outputs are the two dominant API patterns for giving large language models (LLMs) a way to control software. Both can power robust agents in 2026, but they optimize for different axes. Function-calling is about letting the model choose and invoke tools with typed arguments. Structured outputs are about constraining the model’s text to a schema so you can deterministically parse, route, or verify. The choice is less about which is “better” and more about the failure modes, cost envelope, and control surface your application needs.

Across thousands of agent deployments we’ve seen a simple pattern: - If the agent must orchestrate many external systems and side effects safely, function-calling (or “tool calling”) is often the backbone, with structured outputs used inside the loop for planning and verification. - If the agent must produce clean data for downstream code (ETL, enrichment, routing, analytics), structured outputs with strong schema constraints and validators win on cost, latency, and debuggability. - Hybrid designs are the norm in 2026: constrained, structured planning → tool selection and calling → structured verification/critique → commit.

This article is an engineering-first, vendor-agnostic map of the trade space with cost models, reliability levers, and implementation blueprints. Where relevant we’ll name today’s concrete tools: OpenAI/Anthropic tool-calling, Gemini function calling, JSON Schema, Pydantic, Guardrails, Outlines and grammar-constrained decoding, LangChain and LlamaIndex agents, Temporal/Dagster/Argo for orchestration, and evaluators you can run in CI.

If you want a guided build path that combines these techniques in production, consider the AI application development: LLM APIs, RAG, vector stores, prompt engineering, agentic workflows program from Refonte Learning.

The two API patterns in brief

Function-calling (aka tool-calling)

Function-calling adds a discrete decision step: the model returns a tool name and a typed argument object. The host system then invokes that tool (HTTP service, database query, vector search, Python function) and returns results to the model. The model may iterate—choose another tool, or produce a final answer.

Core characteristics: - Selection: model chooses among N tools (dynamic or fixed registry). - Arguments: JSON-like structure, often validated against tool signatures. - Turn-taking: stepwise loop; each iteration adds tokens and latency. - Side effects: tools can mutate state; host must enforce safety, idempotency, and observability.

Where it shines: - Multi-step workflows (plan → act → observe → revise). - Integrations with SaaS and infra (ticketing, billing, Kubernetes, Snowflake). - When human-in-the-loop approvals and audits are required.

Structured outputs

Structured outputs constrain the model to emit data that matches a schema (e.g., JSON Schema, Pydantic model). You get deterministic parsing and can wire the result into downstream code without brittle regexes.

Core characteristics: - Single decision: generate output that validates against schema. - Determinism: grammar/decoder constraints reduce format drift. - Cost and latency: usually cheaper/faster than iterative tool loops.

Where it shines: - Extraction, classification, and routing. - Planning objects (task lists, SQL queries, intents) consumed by non-LLM code. - Batch data ops where throughput and cost dominate.

Vocabulary alignment for 2026

  • Tool/function calling: vendor terms differ but semantics align. OpenAI “tool calls,” Anthropic “tool use,” Google “function calling,” and open-source libraries in LangChain/LlamaIndex provide similar abstractions.
  • Structured output modes: JSON schema, Pydantic, TypeScript-like types, and grammar-constrained decoding (JSON only, regex-like, or full context-free grammars).
  • Constrained decoding: sampler enforces schema during generation (e.g., Outlines, Guidance, llama.cpp grammars). Distinct from “JSON mode” that merely biases the model.
  • Router: policy that decides which model or toolchain to use; can be learned or rule-based.
  • Orchestrator: system that coordinates steps and retries (Temporal, Dagster, Airflow, Argo Workflows).

The core tradeoff: control surface vs. throughput

  • Control surface (function-calling): maximum flexibility, explicit tool selection, mid-flight observation handling, human approvals, rollback, and compensating actions. The cost is extra tokens and latency per turn, plus orchestrator complexity.
  • Throughput (structured outputs): minimum overhead, strong schema guarantees, deterministic parsing. The limit is you cannot observe external state mid-generation without a second pass or a function-call loop.

Most robust agents in 2026 combine both: structure the plan, call tools, structure the postconditions, then render.

Reliability: failure modes and mitigations

Function-calling failure modes

1) Wrong tool selection - Symptom: model picks an email-sending tool instead of a CRM lookup. - Mitigation: distinct, minimal tool descriptions; discriminative routing upfront; penalize generic tool names.

2) Malformed arguments - Symptom: tool call with missing/invalid fields. - Mitigation: strict JSON schema for each tool, automatic argument validation and repair, and few-shot examples of valid calls.

3) Hallucinated entities - Symptom: calling a tool with IDs or fields that do not exist. - Mitigation: pre-validation lookups, referential integrity checks, and rejecting unverifiable calls.

4) Infinite/degenerate loops - Symptom: tool-use thrashing, repeated function calls. - Mitigation: iteration caps, token budget ceilings, watchdog rules, and summarization of state each loop.

5) Side-effect drift - Symptom: partial success with side effects (e.g., created ticket but failed to update status). - Mitigation: idempotency keys, two-phase commit patterns, compensating transactions, and human approval gates.

Structured-output failure modes

1) Format drift - Symptom: extra commentary, trailing commas, or partial JSON. - Mitigation: grammar-constrained decoding; if unavailable, JSON repair with a deterministic parser and validator.

2) Semantic invalidity - Symptom: JSON validates syntactically but violates business rules. - Mitigation: post-generation validators; reflexive critique step that regenerates invalid fields; few-shot counterexamples.

3) Overfitting schema - Symptom: schema too rigid to capture real-world nuance. - Mitigation: versioned schemas, optional fields with enums, and deprecation plans.

4) Hidden coupling - Symptom: downstream services silently rely on model quirks. - Mitigation: contract tests in CI with gold outputs and fuzzing invalid permutations.

Cost model in 2026

You control spend by understanding where tokens and milliseconds go.

Function-calling cost components

  • Prompt context inflation: tool catalogs and prior steps consume input tokens.
  • Tool selection turns: every loop adds output tokens (tool call), input tokens (tool result), and latency.
  • Observation fan-out: tools that return large payloads (full documents, logs) explode tokens.
  • Guardrails: validation/repair cycles add turns.

Cost optimization levers: - Compress tool docs to the minimum discriminative description; use embedding lookup of tool help instead of inlining full catalogs. - Summarize tool responses at the boundary; prefer IDs over raw text. - Cap iterations with policy; require a “stop reason” field and a deterministic termination rule. - Use small models for early selection or schema repair; reserve high-end models for final synthesis.

Structured outputs cost components

  • Single-pass generation: prompt + constrained decoding often yields minimal retries.
  • Validation passes: cheap retries on invalid fields rather than full regeneration.
  • Batching: parallelizable on server or GPU, high throughput per dollar.

Cost optimization levers: - Grammar constraints to avoid repair passes. - Field-wise regeneration instead of whole-object retry. - Lightweight validators (pure Python/TypeScript, WASM) before heavier LLM repair.

Latency considerations

  • Function-calling introduces turn latency: LLM decode + tool call + network + validation. With 2–5 loops, P95 latency can become unacceptable for end-user UIs.
  • Structured outputs usually complete in one pass. With constrained decoding, you may sacrifice a small amount of raw tokens/sec performance for deterministic JSON but still beat multi-turn loops handily.
  • Streaming: for chat UX, stream natural language content, but gate side effects behind a confirmed plan that is structured and verified.

Observability, evaluation, and SLOs

Production agents need SLOs: - Tool-call success rate: valid tool chosen, valid args, tool not failing. - Schema validity rate: percentage of outputs passing syntactic and semantic checks on first attempt. - Token and turn budgets: average/95th per request. - Latency buckets: P50/P95 overall and per-step (decode, tool call, validator). - Safety/guardrail violations: blocked outputs, policy triggers.

Implementation: - Structured logs for each tool call (name, args hash, idempotency key, result metrics). - Persist the plan and the final state; store diffs for audits. - Canary tests: inject synthetic tasks that measure drift. - Evals: curated datasets covering ambiguous intents, adversarial inputs, and rare tool combos.

Security and safety posture

  • Least-privilege tools: narrow scopes, per-environment credentials, and short-lived tokens.
  • Idempotency: client-provided idempotency keys in every mutating tool call.
  • Output filters: PII scrubbing, prompt-leak prevention, and policy checks pre-commit.
  • SSRF/command injection: sanitize any tool that touches networks or shells; whitelist hosts and commands.
  • Human approval: gated steps for financial transfers, user messaging, or infra changes.

Vendor capabilities snapshot in 2026

  • OpenAI: tool calls with JSON arguments, structured output APIs with JSON Schema, parallel tool calls, and function result messages.
  • Anthropic: tool use with signatures, improved constitutional policies, structured output modes with schema hints.
  • Google Gemini: function calling, robust JSON-compliant output options.
  • Open-source stacks: llama.cpp grammar constraints; vLLM/XTuring integrations; Outlines and Guidance for CFG-based decoding; Jsonformer-like decoders; Guardrails for RAIL/JSON Schema enforcement.
  • Orchestration: Temporal, Dagster, Argo Workflows; serverless runtimes like Modal or AWS Lambda for step isolation.

Your architecture should not overfit to a single provider; keep schemas and tool registries provider-agnostic.

Design patterns that work in practice

1) Plan-then-act (structured plan → function calls)

  • Step A: Generate a structured plan object: intents, constraints, risk flags, and required tools.
  • Step B: Loop through plan steps, each mapping to one tool call; validate after each action.
  • Step C: Generate a structured postcondition summary and a human-readable message.

Benefits: human-auditable plan, tighter cost control, deterministic exit.

2) Critic-sandbox-commit

  • Draft: structured output plan with cost/risk estimates.
  • Sandbox: use function calls restricted to read-only tools to validate assumptions.
  • Critique: second LLM pass revises plan, referencing sandbox evidence.
  • Commit: execute mutating calls with idempotency and approvals.

Benefits: reduces costly rollbacks and improves safety on high-stakes tools.

3) Router-first minimal tools

  • Route with a small model to a sub-agent with 3–5 purpose-built tools.
  • Within each sub-agent, prefer structured outputs for data shape; call tools only when the plan requires external state.

Benefits: less tool confusion, lower token bloat.

4) Structured extraction + deterministic executor

  • Use structured outputs to extract parameters and constraints.
  • Feed into a deterministic non-LLM executor (SQL runner, dbt task, ArgoCD deployment) with retries and monitors.

Benefits: maximum determinism for infrastructure and data pipelines.

Tool registries and argument schemas

  • Names: tool names should be specific and disjoint (e.g., create_jira_issue vs. email_customer_smtp), with 1–2 line descriptions focusing on disambiguating features.
  • Signatures: specify types, enums, required vs optional, and length bounds.
  • Examples: include a single canonical example per tool; avoid flooding the prompt with dozens of variants.
  • Dynamic registry: fetch available tools per request (tenant- or role-based), rather than dumping a global catalog.

For implementations, Pydantic, JSON Schema, and TypeScript types are common. Generate both runtime validators and documentation from a single source of truth.

Constrained decoding vs. “JSON mode”

  • JSON mode nudges the model; it can still break formatting under pressure.
  • Grammar-constrained decoding (CFG) eliminates formatting errors by construction. If the token is not allowed by the grammar, it will not be emitted.
  • In 2026, use grammar constraints for high-throughput extractors and routers. For open-ended content, pair a loose schema with post-validators.

Measuring and improving reliability

  • Golden sets: curated tasks with expected tool sequences and outputs.
  • Perturbation tests: randomize field order, insert adversarial whitespace, or vary irrelevant details to detect brittle prompts.
  • Contract tests: treat tool signatures and output schemas as APIs; lock them with versioned tests.
  • Health dashboards: live rates of invalid calls, schema repair attempts, and blocked side effects.

Cost envelopes by use case

  • CRUD agents on SaaS/infra: function-calling with 1–3 steps per task; higher latency but acceptable for back-office automation. Budget for retries and approval steps.
  • ETL/enrichment: structured outputs dominate. Batch-friendly, grammar-constrained decoding, field-wise retries.
  • Chat copilots: hybrid—structured intents and tool selection, stream NL content, gate side effects after plan approval.
  • Analytics and BI: structured generation of SQL with schema-aware validators; function calls for lineage and permissions checks.

Example architectures

A) Support triage copilot

  • Step 1: Structured classification: intent, priority, product area, PII flags.
  • Step 2: If confidence < threshold, escalate to human.
  • Step 3: If automation allowed, function-call to create or update tickets, add tags, and draft response.
  • Step 4: Postcondition summary structured for dashboards.

Key considerations: per-tenant tool registry, audit fields (who/when/why), and cost ceilings.

B) Data pipeline assistant

  • Structured output: dbt node, SQL transformation, test assertions, run-time window.
  • Deterministic executor: run in staging; capture test failures.
  • Critique pass: LLM reviews staging results (read-only tool) and proposes edits.
  • Commit gated by approval.

Key considerations: schema linting, column-level lineage, and security boundaries.

C) Cloud ops agent

  • Structured plan: desired state (Kubernetes manifests, HPA settings), risk level, rollback strategy.
  • Sandbox: read-only describe/list calls; drift detection.
  • Function calls: apply changes via ArgoCD/Flux with idempotency.
  • Postconditions: structured roll-forward/rollback instructions.

Key considerations: blast-radius controls, SLO-aware changes, and policy guardrails.

Practical prompts and schemas

Structured planning schema (example)

  • fields: goal, constraints[], required_tools[], stop_reason, risk_level{low|med|high}
  • validators: every required tool must map to an allowed registry item; risk-level high forces human approval.

Tool argument schema (example)

  • create_invoice(customer_id: string, line_items: array of {sku, qty, unit_price}, currency: enum[USD,EUR,...])
  • soft limits: max 50 line items, currency required.
  • validation path: schema → business rules (tax rules per region) → idempotency key derived from sorted payload.

Postcondition summary schema

  • fields: actions_taken[], external_ids[], success:boolean, residual_risks[], next_steps[]

These schemas keep the agent honest and auditable.

When to choose structured outputs first

  • Data-in, data-out: no side effects, just shaping content for downstream systems.
  • High throughput targets: millions of rows or documents per day.
  • Strict contracts: APIs that will break if the model free-forms.
  • Cost-sensitive: every extra turn magnifies spend; constrained decoding drastically reduces invalidation costs.

When to choose function-calling first

  • Real-world actions: you must inspect external state and act in stages.
  • Uncertain paths: tool choice depends on observations not known at prompt time.
  • Compliance: need approvals, audit trails, and compensating actions.

The hybrid default in 2026

Most teams deploy a hybrid because it cleanly separates concerns: - Structure: generate a verifiable plan object; cheap. - Act: function-call minimally; expensive but necessary. - Verify: structured postconditions to detect drift; cheap. - Render: human-readable content, optionally streaming.

This yields better reliability, clearer logs, and controllable costs.

Guardrails, validators, and repair loops

  • Syntax guardrails: grammar constraints and strict JSON schemas.
  • Semantic validators: business rules in code; don’t push logic into prompts when a function can check it deterministically.
  • Repair strategy: prefer targeted regeneration—ask the model only for fields that failed validation, including validator error messages.
  • Safety rails: LLM-based policies (e.g., refusal when risk level high) should be backed by deterministic checks to prevent bypass.

Testing and CI

  • Unit prompts: small, deterministic checks on schemas and tool arguments.
  • End-to-end golden flows: orchestration-level tests executing mock tools and verifying postconditions.
  • Drift monitors: track schema validity over time; on drift, freeze deploys and run regression suites.
  • Replay harness: store anonymized traces; replay against new models before rollout.

Orchestration and state

  • Temporal/Dagster for compensating transactions and retries.
  • State model: per-request state includes plan, observations, tool call results, approvals, and postconditions. Serialize this state for auditable recovery.
  • Idempotency: derive keys from stable, sorted payloads and use them across retries; tools must treat duplicates as no-ops.

Data handling and privacy

  • Minimize PII in prompts; pass only IDs when possible and fetch details with tools server-side.
  • Redaction: mask PII in logs while keeping hashed references for correlation.
  • Residency: tools should be region-aware; route requests to comply with data localization.

Selecting models for each step

  • Planning and classification: small-to-mid models with structured outputs.
  • Tool selection and reasoning: mid-to-large models depending on ambiguity.
  • Long-form rendering: large models with style controls.
  • Repair/validation: smallest effective model; deterministic validators first.

This mixture optimizes cost without sacrificing reliability.

Monitoring economics: tokens, turns, and tails

  • Tokens: watch average and tail growth as prompts and tool catalogs evolve.
  • Turns: cap max turns; alert on degenerate loops and retries beyond 2.
  • Tails: P95/P99 latency spikes often correlate with large tool payloads; summarize aggressively.

Migration playbook: from free-form chat to robust agents

1) Introduce structured outputs for intents and slots; keep chat responses. 2) Add a single, tightly scoped tool with strict arguments; observe behavior. 3) Expand toolset with a registry and per-tenant availability. 4) Insert a structured postcondition summary and an approval gate. 5) Move validations from prompts into code; reduce repair passes. 6) Add CI evals and replay harness; version schemas.

Developer ergonomics

  • Codegen: derive tool stubs and validators from a single schema source.
  • Trace UX: clickable traces that collapse loops, show diffs, and surface validator failures.
  • Incident response: on-call runbooks for stuck workflows, with safe abort and rollback.

Anti-patterns to avoid

  • One giant tool: a mega-function that does everything induces confusion and wrong selections.
  • Schema sprawl: dozens of near-duplicate schemas create maintenance drift.
  • Prompt-only guarantees: relying on “please return valid JSON” without validators or grammar constraints.
  • Blind trust: executing side effects without idempotency, approvals, or audits.

Example code blueprints (language-agnostic pseudocode)

Structured plan generation

  • input: user_request
  • call LLM with grammar-constrained schema Plan
  • if invalid: regenerate fields that failed
  • persist: plan with version and hash(user_request)

Tool loop

  • for step in plan.steps:
  • validate args via schema
  • dry-run if supported; else, read-only checks
  • if risk > threshold: request approval
  • call tool with idempotency_key
  • record observation summary (bounded size)

Postcondition and render

  • call LLM to produce Postcondition schema
  • if Postcondition.success == false: propose rollback plan (structured)
  • render human message; stream if needed

Procurement checklist for 2026

  • Does the provider support grammar-constrained decoding, not just JSON hints?
  • Can I define tool signatures with JSON Schema and validate arguments automatically?
  • Is parallel tool calling supported and observable?
  • What are the per-step token/latency metrics and billing increments?
  • How do retries and idempotency interact with the platform?
  • Can I export raw traces for external observability and audits?

Cost case study sketch

Scenario: customer onboarding automation. - Baseline: chat-only. Low cost but manual ops. - Structured plan + postconditions: +5–10% token overhead; order-of-magnitude reduction in human escalations for routine tasks. - Function-calling with 2–3 tools (CRM, KYC, billing): +1–2 turns; latency increases but total operator minutes drop sharply. - Optimization: grammar-constrained decoding for plan/postconditions and argument summarization for tool outputs brings costs back near baseline per successful case.

Even without inventing numbers, the qualitative picture holds across teams: structure early and late; use function calls only when the world must be observed or changed.

Team topology and ownership

  • Platform team: owns schemas, tool registry, and orchestration.
  • App teams: own domain tools and validators.
  • Risk/compliance: defines approval policies and audit retention.
  • MLOps: owns evals, model selection, and drift monitors.

Documentation and versioning

  • Semantic version schemas (major.minor.patch); deprecate fields over quarters, not weeks.
  • Changelog entries must include “migration notes” for downstream services.
  • Tool registry docs: one-pager per tool with examples and known failure modes.

Benchmarks that matter

  • First-pass schema validity rate.
  • Valid tool-call rate without manual repair.
  • End-to-end success (task solved without human intervention) at P50 and P95 latencies.
  • Cost per successful task (not per token).

What’s next in 2026–2027

  • Stronger native structured output guarantees across major providers.
  • Tighter integration between function calling and vector search, enabling context-aware tool descriptions without dumping catalogs.
  • Better field-wise regeneration APIs and partial decoding constraints.
  • Wider adoption of deterministic executors for side effects with LLMs focusing on planning and explanation.

How to decide quickly

  • If your agent primarily shapes data: choose structured outputs first, with grammar constraints and validators. Add function calls only to fetch missing facts.
  • If your agent primarily changes the world: choose function-calling with strict schemas, idempotency, and an orchestrator. Use structured outputs for plans and postconditions.
  • Default hybrid: structured plan → minimal tool loop → structured verification → render.

For practitioners who want to build this muscle with feedback and real projects, explore Refonte Learning’s AI developer program. It blends architecture patterns, reliability engineering, and cost control in capstone builds.

About Refonte Learning

Refonte Learning is a practitioner-led EdTech platform focused on AI, data, cloud, DevOps, and software engineering. Our instructors ship production systems and bring hard-won patterns into the classroom. If you want hands-on training in agentic application design, our study-and-internship model pairs structured learning with real-world builds so you can become an AI application engineer who ships reliable systems.

Bottom line

Function-calling and structured outputs are complementary tools for agent design in 2026. Use structured outputs anywhere you can enforce contracts cheaply; use function-calling where the agent must perceive and act in the world, with tight safety and observability. Most winning systems are hybrids that plan and verify with structure and act with minimal, auditable tool calls. Build around schemas, validators, and idempotent tools, and your agents will be cheaper, faster, and more reliable to operate at scale.

Finally, if you’re assembling a roadmap or hiring plan around these capabilities, consider the study-and-internship path for AI builders from Refonte Learning for a structured way to de-risk your first production agents.