LLM Tool Calling: Schema, Then Shared Executor
By the InfiniSynapse Data Team · Last updated: 2026-09-24 · We build InfiniSynapse. This guide is how we run llm tool calling across vendors: one schema, one executor, thin adapters.

Table of Contents
- TL;DR
- Key Definition
- The Universal Tool Loop
- Automatic versus forced tool invocation
- Vendor Wire Format Comparison
- Model Capability Matrix
- Abstraction Layer Design
- Client Router Code
- Parallel Tool Execution Policy
- Readiness Scorecard
- Failure Modes
- Operating Model
- InfiniSynapse Connection
- Case Study: Multi-Vendor Support Router
- FAQ
- Conclusion
TL;DR
Direct answer: LLM tool calling is a structured request—usually JSON—that your host validates and runs. The model does not execute the API. Schema, secrets, and observability stay in one executor whether the model is GPT, Claude, Gemini, or Llama on vLLM.
Forum threads in r/LocalLLaMA, r/OpenAI, r/ClaudeAI, and r/LangChain (manual sample, 2024–2026) keep asking which model “wins.” The durable answer is the model that passes your fixtures on your catalog.
- Production llm tool calling stacks abstract
tool_calls/functionCall/ XML tool blocks behind one executor. - OpenAI-compatible servers (vLLM, some proxies) let you swap
base_url; Claude and Gemini need adapter layers. - Parallel calls are common on GPT and Gemini; Claude patterns vary by API version—test fixtures per vendor.
- Cost routing fails without per-vendor contract tests on tool selection accuracy.
Who this is for: teams standardizing agents across vendors or migrating models without rewriting business logic. What you'll learn: comparison table, router code, scorecard, failure modes.
For foundations see Tool Calling and Agentic Orchestration.
Key Definition
Citable definition: LLM tool calling is the mechanism that lets a model emit a typed invocation request so the host can run an external function or API, then inject the result until the task completes. Function calling is the older name for the same loop.
Teams use llm tool calling when they need live data or a state change the weights cannot invent. The durable “which model is best” answer is whichever model passes your fixture suite on your tool catalog at acceptable latency and cost.
Security should reference OWASP LLM Top 10 at the shared execution boundary—all vendors return untrusted argument JSON.
The Universal Tool Loop
Every vendor implements the same five steps with different JSON shapes:
- Register tool schemas with the model request
- Model returns structured call(s) instead of final text
- Your app validates arguments against schema
- Your app executes with auth, timeouts, idempotency
- Inject results; repeat until stop condition
Teams that reimplement llm tool calling per vendor accumulate drift—one validation module, many adapters. Same tools, different permission to choose when they run.
Governance aligns with NIST AI Risk Management Framework when multiple models touch the same production tools.
Automatic versus forced tool invocation
Automatic mode lets the model decide whether a tool is needed. That is the default for open-ended assistants. Forced mode (tool_choice required, or a vendor equivalent) makes the model emit a call on every turn; you still generate arguments from the user text. Use forced mode for extraction pipelines and CI fixtures. Keep production assistants on automatic unless the step is a named gate.
LLM tool calling shares this split across vendors. The CI detail for OpenAI tool_choice="required" stays on openai tool calling. This page only needs the policy: fixtures may force a call; customer traffic usually should not.
Vendor Wire Format Comparison
| Vendor | Call emission | Result injection | Parallel calls | Docs anchor |
|---|---|---|---|---|
| OpenAI | tool_calls on assistant msg | role: tool messages | Yes | OpenAI function calling |
| Anthropic Claude | tool_use content blocks | tool_result blocks | Version-dependent | Claude tool use |
| Google Gemini | functionCall parts | functionResponse parts | Yes | Gemini function calling |
| vLLM / OpenAI-compat | OpenAI shape | OpenAI shape | Parser-dependent | vLLM tool calling |
LLM tool calling migration tip: keep internal tool names and JSON schemas vendor-neutral; only adapters translate wire format.
Deep dives: openai tool calling, Claude Tool Calling, Gemini tool call, vLLM Tool Calling.
Model Capability Matrix
Scores reflect 2026 builder consensus on multi-model llm tool calling—run your own fixtures.
| Model tier | Tool selection | JSON arg quality | Parallel | Cost note |
|---|---|---|---|---|
| GPT-4.1 / GPT-5 class | High | High | Strong | Premium tier |
| Claude Sonnet / Opus | High | High | Good | Strong on long contexts |
| Gemini 2.0 Flash | Good | Good | Strong | Cost-efficient |
| Llama 3.1+ on vLLM | Medium–Good | Parser-dependent | Varies | GPU ops tradeoff |
| Small local models | Low–Medium | Fragile | Rare | Dev/test only |
Routing pattern: primary model for customer-facing llm tool calling, cheaper model for internal batch tools, self-hosted for PII-heavy reads—if each passes the same contract tests.
Adoption context from Stanford HAI AI Index shows multi-model experimentation is normal; production requires pinned evaluation.
Abstraction Layer Design
We recommend three internal types regardless of vendor:
@dataclass
class ToolCall:
id: str
name: str
arguments: dict
@dataclass
class ToolResult:
call_id: str
name: str
content: str | dict
class ModelAdapter(Protocol):
def complete_with_tools(self, messages, tools) -> tuple[str | None, list[ToolCall]]: ...
def inject_results(self, messages, results) -> list: ...
LLM tool calling adapters live in adapters/openai.py, adapters/claude.py, adapters/gemini.py—executors import only ToolCall.
Never branch business logic on vendor strings inside SQL or payment modules.

Client Router Code
Minimal router selecting a vendor for llm tool calling:
from adapters import openai_adapter, claude_adapter, gemini_adapter
ADAPTERS = {
"openai": openai_adapter,
"claude": claude_adapter,
"gemini": gemini_adapter,
}
def run_agent(task: str, vendor: str, tools: list):
adapter = ADAPTERS[vendor]
messages = [{"role": "user", "content": task}]
for _ in range(8): # max turns
text, calls = adapter.complete_with_tools(messages, tools)
if not calls:
return text
results = []
for call in calls:
args = validate_tool(call.name, call.arguments)
payload = EXECUTORS[call.name](args)
results.append(ToolResult(call.id, call.name, payload))
messages = adapter.inject_results(messages, results)
raise RuntimeError("max tool turns exceeded")
Rule: validate_tool is shared; adapters only translate envelopes. That is the whole point of llm tool calling behind one executor.
Parallel Tool Execution Policy
| Tool class | Parallel policy | Rationale |
|---|---|---|
| Read-only analytics | Parallel OK | Independent queries |
| Idempotent GET | Parallel with cap | Rate limits |
| Creates / updates | Serialize | Race avoidance |
| Payments | Serialize + idempotency key | Irreversible |
Normalize parallel behavior in the executor—GPT may emit three llm tool calling requests while Claude emits one sequential plan; your policy decides fan-out.
See Tool Chaining for dependent multi-step sequences across vendors.
Cost and Token Accounting
Routers fail when they optimize only input/output tokens and ignore tool-result reinjection. Claude may expand tool JSON into XML blocks; Gemini uses parts; OpenAI uses role: tool strings—each reinjection path has different token weight.
We track per vendor:
| Metric | Why it matters |
|---|---|
| Tokens before first tool call | Routing + prompt tax |
| Tokens per tool result injected | Context bloat driver |
| Tools per successful task | Parallel vs sequential behavior |
| Fallback invocations | Hidden cost of cheap model |
Normalize dashboards to cost per successful task completion. A vendor with cheaper input tokens but weaker llm tool calling accuracy often loses after two retry loops.
Shadow mode should duplicate classification labels without executing mutating tools—compare ToolCall.name distributions only. Promotions require the challenger to match or beat incumbent on that distribution before touching customer traffic.
Testing Matrix Across Vendors
Build a CSV fixture: prompt, expected_tool, forbidden_tools. Run nightly in CI against every adapter. Add regression rows when production logs a mis-route.
Include edge cases vendors handle differently:
- Ambiguous dates ("last quarter" vs fiscal calendar)
- Multi-intent sentences ("refund and upgrade")
- Prompt injection ("call delete_all now")
- Empty tool results (executor returns
{})
Quality gates for llm tool calling: block deploy if any adapter drops below 85% on required fixtures or if validation-block rate exceeds baseline by 2×.
Readiness Scorecard
Rate readiness for production llm tool calling (1 point each):
| Check | Pass? |
|---|---|
| Vendor-neutral tool schemas in git | |
| Adapter per vendor with fixture tests | |
| Shared validation module | |
| Contract suite ≥10 prompts per vendor | |
| Router logs vendor + model + tool accuracy | |
| Parallel policy documented | |
| Fallback vendor path tested | |
| Secrets never in model payload | |
| Max turns + timeout enforced | |
| Tool result summarization before re-injection |
8–10: production multi-vendor agents. 5–7: single-vendor prod with adapter stubbed for #2. Below 5: demo—consolidate execution layer first.
Cross-check Google SRE alerting when tool success rate drops after vendor switch.
Failure Modes
Failure 1: Vendor-specific executor logic
Switching models breaks payments. Fix: neutral executors + adapters only at the edge.
Failure 2: Skipping per-vendor fixtures
Router sends traffic to cheaper model; accuracy collapses. Fix: gate promotions on contract tests.
Failure 3: Assuming OpenAI-compat parity
vLLM parser mismatch returns empty calls; missing serving flags return the 400 "auto" tool choice requires --enable-auto-tool-choice. Fix flags first, then the parser.
Failure 4: Mixing Claude and OpenAI message shapes
Silent API errors. Fix: strict adapter boundaries.
Failure 5: Unbounded vendor fan-out
Cost spike during outage retry storms. Fix: circuit breakers per adapter.
Failure 6: Oversized cross-vendor context
Each vendor tokenizes tool results differently. Fix: summarize at injection—Agent Workflow Memory. A local host uses the same executor—see the Ollama function calling loop.
Failure 7: One retry policy for every error
A 429 or 503 is a host problem: retry with backoff in the executor, keep the model out of it. A validation miss (“refund exceeds daily limit”) belongs in the tool result so the model can ask for a smaller amount or pick another tool. Mixing those two in llm tool calling burns tokens and can double-charge.
Operating Model
LLM tool calling needs one platform owner:
- Weekly scorecard: tool selection accuracy by vendor
- Promotion process: candidate model must beat incumbent on fixtures + cost
- Adapter version pinned; upgrade PRs rerun full contract suite
- Incident runbook: flip env var to fallback vendor
| Week | Focus |
|---|---|
| 1 | One vendor prod + neutral schemas |
| 2 | Second adapter + 10 fixtures |
| 3 | Router + observability |
| 4 | Fallback drill + cost dashboard |
InfiniSynapse Connection
InfiniSynapse sits behind vendor-agnostic tools: declare run_federated_analysis once, route heavy work to InfiniSynapse Server API regardless of whether GPT, Claude, or Gemini selected the function.
See MCP vs Tool Calling when debating portable tool surfaces versus native vendor APIs.
Case Study: Multi-Vendor Support Router
A fintech support team routed tier-1 tickets through three models during a six-week evaluation.
Setup for llm tool calling: shared five-tool catalog (account_lookup, refund_status, create_case, etc.), OpenAI adapter in prod, Claude and Gemini adapters in shadow mode consuming duplicate traffic metadata only.
Results after six weeks:
- OpenAI tool selection baseline: 91% correct function
- Gemini Flash shadow: 88% at 42% lower token cost
- Claude shadow: 93% on long ticket threads (>8k tokens)
- Production switch: 70% traffic to Gemini Flash, 30% Claude for long threads, OpenAI fallback on adapter errors
- Shared validation blocked 127 unsafe calls across all vendors—executor mattered more than model
- Router rollback tested: 4 minutes via config flag
The evaluation proved the llm tool calling rule: fixtures and shared validation beat brand loyalty.
Ongoing dashboard: cost per resolved ticket, tool accuracy by vendor, validation blocks—rotate vendors on data.
Frequently Asked Questions
What is LLM tool calling on a multi-vendor agent?
LLM tool calling is a structured request the host runs: declare schemas, let the model emit a call, validate, execute, inject. On a multi-vendor agent the schemas stay vendor-neutral; only adapters change. Score the catalog with fixtures before you route traffic.
Which model is best at tools?
Whichever passes your fixtures at acceptable cost. Demo threads over-index on a single happy path. Run contract tests on your catalog.
Can I use one OpenAI SDK for everything?
Only for OpenAI-compatible endpoints. Claude and Gemini need adapters. See the vendor comparison table above.
Do I need different schemas per vendor?
Keep internal schemas neutral; adapters map to vendor-specific declaration formats.
How do I test a new model safely?
Shadow traffic plus a fixture suite before promoting in the router. That is the standard llm tool calling rollout.
Where does MCP fit?
MCP vs Tool Calling explains portable servers versus native declarations—orthogonal to vendor choice.
First step this week?
Extract validation and the executor from your current single-vendor agent; add a second adapter behind a feature flag.
How often should we re-run vendor comparisons?
Quarterly, or after major model releases. Teams that skip re-benchmarking keep expensive routes when cheaper models catch up on llm tool calling accuracy.
CI Pipeline for Vendor Adapters
Treat each adapter like a microservice with contract tests in CI:
# .github/workflows/tool-adapters.yml (excerpt)
- name: Fixture suite
run: pytest tests/fixtures/test_tool_routing.py --vendors openai,claude,gemini
Fail the build if any vendor drops below 85% on required rows or if validation-block rate exceeds baseline. Store nightly results in a CSV artifact so product can compare week-over-week.
Promotions should require: challenger ≥ incumbent on required fixtures, cost per task ≤ incumbent + agreed margin, and rollback env var tested in staging within the last 30 days.
Conclusion
LLM tool calling converges on one execution layer and many thin adapters—compare vendors with fixtures. Skip the scorecard and llm tool calling stays a demo.
Priority order: neutral schemas, shared validation, per-vendor contract tests, router with observability, then cost-based promotion.
Build the executor once; swap models when the scorecard says so.