OpenAI Tool Calling: Wire Format and the Loop

By the InfiniSynapse Data Team · Last updated: 2026-09-27 · We build InfiniSynapse and write these notes like a builder posting after a Reddit thread—not a brochure for vibe-coded products moving to real APIs and data infrastructure.

Hero image for openai-tool-calling


Table of Contents

  1. TL;DR
  2. Key Definition
  3. OpenAI Wire Format
  4. Chat Completions vs Responses API
  5. Parallel Tool Calls
  6. function_call and tool_calls
  7. Strict Mode and JSON Schema
  8. Execution Loop Code
  9. When a tool fails
  10. Streaming Tool Calls
  11. Readiness Scorecard
  12. Failure Modes
  13. InfiniSynapse Connection
  14. Case Study
  15. Frequently Asked Questions
  16. Conclusion

TL;DR

Direct answer: OpenAI tool calling returns tool_calls: a function name plus JSON-string arguments. Your app validates, runs the function, and appends a role: tool message. The model does not execute the function. Reliability still depends on parallel calls in one assistant turn, and on Chat Completions versus the Responses API when you need server-side state.

If you have spent time in r/OpenAI, r/LocalLLaMA, r/LangChain, and r/vibecoding, you have seen these arguments. Here is what held up when teams wired OpenAI models to real product actions—not the "just pass functions" hype.

  • openai tool calling pattern: register tools → model emits tool_calls → validate args → execute → append role: tool messages → re-infer.
  • Parallel tool_calls in a single response must execute concurrently—sequential handling doubles latency on lookup-heavy tasks.
  • Responses API suits greenfield agents with built-in conversation state; Chat Completions stays the default when you own the message list.
  • Strict JSON Schema mode reduces argument hallucination on enums and nested objects.

Who this is for: engineers standardizing on OpenAI for agent backends. What you'll learn: wire format, API choice, parallel execution, code, scorecard, failure modes.

For cross-vendor context see Tool Calling and LLM Tool Calling.

Key Definition

Key Definition: openai tool calling covers how OpenAI models emit structured tool_calls—function name plus JSON-string arguments—that your runtime validates, executes, and returns via tool role messages for multi-step agent behavior.

openai tool calling matters when your GPT integration returns polished text but never touches your database, billing API, or warehouse with auditable side effects.

OpenAI documents the canonical format in OpenAI function calling guides. Anthropic and Gemini differ on the wire—see Claude Tool Calling and Gemini tool call—but the execution loop you build stays the same.

Security should reference OWASP LLM Top 10—especially excessive agency when write tools lack approval gates.

OpenAI Wire Format

Every openai tool calling integration starts with tool registration in the request:

{
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "lookup_order",
        "description": "Fetch order status by order_id. Use only for existing orders; never invent IDs.",
        "parameters": {
          "type": "object",
          "properties": {
            "order_id": { "type": "string", "description": "UUID from checkout confirmation." }
          },
          "required": ["order_id"]
        }
      }
    }
  ],
  "tool_choice": "auto"
}

When the model decides to act, the assistant message includes tool_calls:

{
  "role": "assistant",
  "content": null,
  "tool_calls": [
    {
      "id": "call_abc123",
      "type": "function",
      "function": {
        "name": "lookup_order",
        "arguments": "{\"order_id\": \"ord_9f2a\"}"
      }
    }
  ]
}

Your runtime parses arguments as JSON, executes, then appends:

{
  "role": "tool",
  "tool_call_id": "call_abc123",
  "content": "{\"status\": \"shipped\", \"carrier\": \"UPS\"}"
}

openai tool calling reviewers flag two mistakes immediately: treating arguments as an object (it is a string on the wire) and omitting tool_call_id on the result message—which breaks the model's ability to correlate results.

Set tool_choice to "none" (do not call a tool), "auto" (default; the model decides), "required" (must call some tool), or {"type": "function", "function": {"name": "lookup_order"}} to force one name. Production stays on "auto". CI can use "required".

function_call and tool_calls

OpenAI's June 2023 API used a functions array and a single function_call field. In November 2023 that surface became tools and tool_calls. OpenAI tool calling today registers each tool with "type": "function". The assistant message returns an array. Each item has id, type, and function.name, plus function.arguments as a string.

Map an old function_call onto one tool_calls entry only while a legacy client still sends it. New code reads tool_calls and echoes tool_call_id on the result. A payload with function_call: null beside a filled tool_calls array is the current shape, not a missed call.

Chat Completions vs Responses API

OpenAI now offers two surfaces for openai tool calling workloads:

SurfaceYou manageBest for
Chat Completions (/v1/chat/completions)Full message array, tool results, historyExisting apps, custom memory, multi-vendor adapters
Responses API (/v1/responses)Prompt + prior response IDs; server stores stateNew agents, simpler client, built-in tool loop helpers

Chat Completions remains the workhorse: you append assistant tool_calls and tool results to messages, then call again until finish_reason is stop with no pending tools. Full control over truncation, compression, and cross-model portability.

Responses API reduces boilerplate for tool loops—the API can chain tool execution rounds when configured with tools and appropriate store settings. Trade-off: tighter coupling to OpenAI's state model; migrating off requires exporting conversation history.

Practical openai tool calling guidance: greenfield OpenAI-only agents can prototype on Responses API; production systems with existing message stores, LangChain graphs, or multi-provider fallbacks should stay on Chat Completions until Responses API parity is proven for your observability stack.

See OpenAI Responses API documentation for current tool-loop semantics—do not assume Chat Completions behavior maps one-to-one.

Parallel Tool Calls

Recent OpenAI models often emit multiple tool_calls in one assistant turn when lookups are independent—e.g., fetch customer profile and open support ticket in parallel.

openai tool calling anti-pattern: iterating tool_calls sequentially when neither depends on the other.

import asyncio
import json

async def handle_assistant_turn(assistant_message, execute_fn):
    if not assistant_message.get("tool_calls"):
        return assistant_message.get("content")

    async def run_one(tc):
        name = tc["function"]["name"]
        args = json.loads(tc["function"]["arguments"])
        result = await execute_fn(name, args)
        return {
            "role": "tool",
            "tool_call_id": tc["id"],
            "content": json.dumps(result),
        }

    tool_results = await asyncio.gather(
        *[run_one(tc) for tc in assistant_message["tool_calls"]],
        return_exceptions=True,
    )
    return tool_results

Wrap asyncio.gather with per-tool timeouts. If one parallel call fails, return structured error in that tool's message so the model can retry or replan—never fail the entire turn silently.

Log parallel_tool_count per request; spikes above three often indicate overlapping tool descriptions—fix schemas before blaming the model.

Independent parallel calls differ from tool chaining—see Tool Chaining when step B requires step A's output.

Strict Mode and JSON Schema

OpenAI supports strict: true on function parameters (where the model tier allows it), enforcing JSON Schema conformance on emitted arguments. For openai tool calling teams fighting enum drift and missing required fields, strict mode is often cheaper than prompt engineering.

Rules that matter in production:

  • Mark every required field in required array—models treat optional fields as skippable under load.
  • Use enum for small fixed sets (status: ["open", "closed"]) instead of free-text descriptions.
  • Avoid additionalProperties: true on objects holding user-influenced data—prompt injection can smuggle keys.
  • Version tool schemas in git; tag tools@v3 in logs when debugging regressions.

Validate arguments server-side even with strict mode—defense in depth against API changes and proxy bugs.

Execution Loop Code

A production openai tool calling turn is five steps: register tools, read tool_calls, run each call, append role: tool, then stop or call the model again.

TOOL_HANDLERS = {
    "lookup_order": lookup_order,
}

def dispatch(name, args):
    handler = TOOL_HANDLERS.get(name)
    if handler is None:
        return {"error": "unknown_tool", "name": name}
    return handler(**args)

Name-based if chains break as the registry grows. Resolve the function from the table, then run it.

MAX_TOOL_ROUNDS = 15

async def run_openai_agent(client, messages, tools, execute_fn):
    for _ in range(MAX_TOOL_ROUNDS):
        response = await client.chat.completions.create(
            model="gpt-4o",
            messages=messages,
            tools=tools,
            tool_choice="auto",
        )
        msg = response.choices[0].message
        if not msg.tool_calls:
            return msg.content

        messages.append(msg.model_dump())
        tool_messages = await handle_assistant_turn(msg.model_dump(), execute_fn)
        messages.extend(tool_messages)

    raise RuntimeError("MAX_TOOL_ROUNDS exceeded")

Credentials stay outside model context. Return { "error": "invalid_arguments", "details": [...] } on schema failure—never raw stack traces. Cap total tool invocations per user session; openai tool calling demos that skip budgets loop until timeout.

Reliability practices from Google SRE apply: log tool name, latency, error code, and sanitized args on every execute.

When a tool fails

Return the error as that call's tool message. An openai tool calling runtime that drops the whole turn gives the model nothing to correct. Say so in the system prompt: on a tool error, retry once with adjusted arguments before answering the user.

MAX_TOOL_ROUNDS stops one request from spinning. A session budget does the same job across turns. Log the tool name, latency, and error code. Do not log secrets or raw stack traces. Independent failures inside asyncio.gather stay on that tool's message so the other calls in the same turn can still return.

Tool Choice and Model Tiers

openai tool calling teams should document which models support parallel tools, strict JSON Schema, and vision-side tools in their pinned API version—not assume feature parity across GPT-4o, GPT-4o-mini, and reasoning-focused tiers.

Use tool_choice: "required" in CI contract tests to force at least one invocation per golden prompt. Production stays on "auto" unless you orchestrate deterministic pipelines (e.g., always call extract_entities before write_crm). If that same "auto" payload hits self-hosted vLLM, the 400 "auto" tool choice requires --enable-auto-tool-choice means the serving process is missing both flags—not that the OpenAI SDK payload is wrong.

When mixing text and tools, assistant messages may include both content and tool_calls. Preserve the full assistant message in history; stripping text loses reasoning traces your on-call engineer needs during incidents.

Reasoning models may emit internal chain-of-thought separately from tool calls—follow current OpenAI guidance on what to log vs redact. openai tool calling observability should always capture tool name, args hash, latency, and status regardless of model family.

Debugging OpenAI Tool Selection

When logs show wrong tool picks, openai tool calling teams fix schemas before swapping models:

  1. Export last 30 mispredictions with user message + tool list version
  2. Cluster by confused tool pairs—usually overlapping descriptions
  3. Add "Do NOT use when…" to the loser tool in each pair
  4. Re-run golden prompts in CI with tool_choice: "required"

Wrong-tool rate above ~10% after description pass usually means too many tools in one request—split into dynamic phases per AI-native orchestration.

Rollout Workflow

Recommended openai tool calling sequence:

StepAction
1One read-only tool on Chat Completions
2Wire-format tests: parse args string, echo tool_call_id
3Add parallel handler with concurrency limit
4Fault injection: timeout, 500, malformed args
5Evaluate Responses API in staging if state ownership fits
6Second tool only after week-one error rate stable

Skipping wire-format tests is how openai tool calling ports from Claude break silently on first multi-tool turn.

Streaming Tool Calls

Streaming openai tool calling integrations receive tool_calls deltas across chunks—index, id, and function.arguments may arrive incrementally.

Buffer until each tool_calls[i] has complete arguments JSON before execute. Partial-parse crashes are a common production bug when teams stream to UI but execute on first chunk.

Pattern: accumulate deltas in a dict keyed by index; on finish_reason: tool_calls, validate and run. For UX, show "Calling lookup_order…" when function.name stabilizes; inject results only after full args parse.

Streaming does not change parallel semantics—still gather independent calls after the assistant turn completes.

Readiness Scorecard

Rate openai tool calling readiness (1 point each):

CheckPass?
Tools registered with when/when-not descriptions
arguments parsed from JSON string, not assumed object
Every tool result includes matching tool_call_id
Parallel tool_calls execute concurrently
Args validated pre-execute; structured errors returned
MAX_TOOL_ROUNDS and per-tool timeouts enforced
Secrets only in runtime, never in prompts
Strict schema or manual validation for enums
Streaming buffers complete args before execute
Chat Completions vs Responses API choice documented

8–10: production beta. 5–7: pilot one workflow. Below 5: fix wire-format bugs before adding tools.

Failure Modes

Failure 1: Missing tool_call_id — model ignores tool results. Fix: echo exact id from assistant turn.

Failure 2: Sequential parallel calls — 2× latency on multi-lookup tasks. Fix: asyncio.gather or thread pool.

Failure 3: Unparsed arguments string — TypeError on execute. Fix: json.loads with try/except → structured error to model.

Failure 4: Wrong API surface — Responses state lost on client refresh. Fix: document state ownership; export history if needed.

Failure 5: Oversized tool results — context blow-up. Fix: truncate to summary + row count; offer follow-up tool.

Failure 6: Mis-cited docs — linking generic AI indexes instead of OpenAI function calling. Fix: cite vendor docs directly.

InfiniSynapse Connection

openai tool calling for data actions: define run_data_analysis in your tools array; proxy submits InfiniSynapse Server API newTask; poll or SSE until complete; inject signed artifact URL as tool content. The model sees one structured result; infra handles long runtime.

Orchestration patterns in Agentic Orchestration, including AI-native orchestration.

Case Study: Inventory Reconciliation

A retail ops team built an openai tool calling agent over GPT-4o with three tools: query_warehouse_skus, query_pos_sales, write_reconciliation_report.

Stack: Chat Completions, Python execution layer, strict schemas on date enums, parallel fetch of warehouse + POS when regions independent, max 14 tool rounds.

Measured pilot (50 SKUs × 12 regions):

  • End-to-end p50: 42s per reconciliation batch
  • SKU match accuracy vs manual sample: 96%
  • False mismatch flags: 8%
  • Runs completing without human fix: 74%
  • Latency saved vs sequential tool handling: ~35% after parallel fix

Wrong-tool rate dropped from 14% to 3% after rewriting descriptions—before any model upgrade. openai tool calling teams often underestimate schema work relative to model selection.

Week-two observability: OpenTelemetry spans per tool; p95 execute 380ms sync, report generation async queue p95 22s.

Frequently Asked Questions

What is OpenAI tool calling?

OpenAI tool calling is the API shape where the model returns tool_calls and your application runs the function. Send the result back as a role: tool message with the same tool_call_id. The model then writes the user-facing answer, or asks for another call.

What is the difference between function_call and tool_calls?

function_call is the older single-call field, paired with a functions array. tool_calls is the current array, paired with tools. Read tool_calls, and send results with tool_call_id.

Conclusion

openai tool calling success is wire-format discipline: parse arguments correctly, correlate tool_call_id, parallelize independent calls, validate before execute, and pick Chat Completions vs Responses API deliberately.

Ship one read-only tool end-to-end with logging before expanding the registry. Compare providers in Claude Tool Calling when adding a second vendor.

OpenAI Tool Calling: Wire Format and the Loop