Gemini Tool Call: Return the Function, Then Run It

By the InfiniSynapse Data Team · Last updated: 2026-09-23 · We build InfiniSynapse and write these notes from production function-calling loops—not a brochure. Product behavior cites Google AI function calling, rechecked 2026-09-23.

Hero image for gemini-tool-calling


Table of Contents

  1. TL;DR
  2. Key Definition
  3. Gemini vs OpenAI Tool Wire Format
  4. AUTO, ANY, and NONE
  5. Send the same id back
  6. AI Studio vs Vertex AI
  7. Function Declaration Design
  8. Client Code with google-genai
  9. Parallel and Sequential Tool Turns
  10. Grounding vs Custom Tools
  11. Architecture Sketch
  12. Readiness Scorecard
  13. Failure Modes
  14. Operating Model
  15. InfiniSynapse Connection
  16. Case Study: Workspace Admin Bot
  17. FAQ
  18. Conclusion

TL;DR

Direct answer: A gemini tool call is JSON your app runs. You declare the function, Gemini returns a functionCall, you execute it, then you send functionResponse with the same id. Gemini does not run the function.

  • Put function_declarations in the tools array. Read functionCall parts on the candidate. Inject functionResponse before the next model turn.
  • AUTO lets the model choose. ANY forces a call. NONE forbids one.
  • Gemini 3 returns an id on every call. Echo that id on the response, not only the function name.
  • Parallel read calls can fan out. A call that needs the previous result stays sequential. Writes stay serialized.
  • Google Search grounding is not your warehouse tool. Mixing a built-in tool with a custom function means you return the thought signature with the function result.

Who this is for: builders wiring Gemini into ops copilots, support bots, or internal data agents. What you'll learn: the call shape, modes, the id round-trip, and the checks when a call never fires.

For general patterns see Tool Calling and Agentic Orchestration.

Key Definition

Key Definition: A gemini tool call is one turn of Google's function-calling loop: declare the tool, receive a functionCall part (name, args, id), execute it in your application, return a functionResponse with that same id, and repeat until the model emits user-facing text.

The call fails open when the declaration is vague, tool_config is NONE, or the previous turn never sent functionResponse. Switching models does not fix that loop.

Secure rollouts should reference OWASP LLM Top 10 at the execution boundary you control after Gemini returns arguments.

Gemini vs OpenAI Tool Wire Format

AspectOpenAI toolsGemini function_declarations
SchemaJSON Schema in parametersOpenAPI-style parameters object
Call shapetool_calls[] on assistant messagefunctionCall part in candidates[0].content.parts
Result injectionrole: tool messagesfunctionResponse part with the same id
Parallel callsMultiple tool_callsMultiple functionCall parts in one turn
Modetool_choicetool_config.function_calling_config.mode

Migrations from OpenAI often fail because developers map tool role messages literally. Gemini expects structured parts, documented in Google AI function calling.

AUTO, ANY, and NONE

Set the mode on tool_config.function_calling_config. Checked against the function calling guide on 2026-09-23.

ModeWhat Gemini does
AUTODefault. The model may return a call or plain text.
ANYThe model must return a function call. allowedFunctionNames limits which names it may pick.
NONEThe model must not return a function call.

allowedFunctionNames belongs with ANY. Setting it under AUTO does not give you a forced call. If a call never fires, read the mode before you rewrite the prompt.

Send the same id back

Gemini 3 returns a unique id on every functionCall. Put that exact id on the functionResponse. Matching only name is not enough when two calls in one turn share a name, or when the next turn has to attach the result to the right request. The function calling guide states the id is always present on Gemini 3.

Governance aligns with NIST AI Risk Management Framework when Gemini tools touch production databases.

AI Studio vs Vertex AI

ConcernGoogle AI StudioVertex AI Gemini
AuthAPI key in envService account + IAM
Data residencyConsumer termsGCP region pinning
VPCPublic endpointPrivate Service Connect options
QuotasDeveloper limitsEnterprise quota + billing
AuditLimitedCloud Logging integration

Production path: prototype in AI Studio, pin the model id, then lift the same declaration JSON to Vertex with Vertex AI function calling for IAM and endpoint URLs.

Compare OpenAI-native patterns in openai tool calling when you run multi-vendor routers.

Function Declaration Design

Gemini chooses among declared functions based on name, description, and parameter property descriptions—same discipline as OpenAI.

Declaration rules we apply in InfiniSynapse pilots:

  • One function per user-visible action (list_workspaces, not do_everything)
  • Enum parameters for bounded choices—reduces hallucinated string args
  • Required fields explicit; optional fields documented in description text
  • Read-only vs mutating tools separated so tool_config can restrict modes

Teams that dump fifty tools in one request see degraded selection. Curate 5–12 active tools per session and swap catalogs by workflow phase.

Client Code with google-genai

Minimal loop with the official Python SDK. Echo function_call.id on the response:

import json
from google import genai
from google.genai import types

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

list_users_decl = types.FunctionDeclaration(
    name="list_active_users",
    description="Return active user count for last N days. Read-only.",
    parameters={
        "type": "object",
        "properties": {
            "days": {"type": "integer", "description": "1-90"},
        },
        "required": ["days"],
    },
)

tools = types.Tool(function_declarations=[list_users_decl])
config = types.GenerateContentConfig(tools=[tools])

response = client.models.generate_content(
    model="gemini-2.0-flash",
    contents="How many active users in the last 7 days?",
    config=config,
)

for part in response.candidates[0].content.parts:
    if part.function_call:
        args = dict(part.function_call.args)
        validate_days(args["days"])  # your guardrail
        result = run_readonly_query(args)
        followup = client.models.generate_content(
            model="gemini-2.0-flash",
            contents=[
                types.Content(role="user", parts=[types.Part(text="How many active users in the last 7 days?")]),
                response.candidates[0].content,
                types.Content(
                    role="user",
                    parts=[
                        types.Part(
                            function_response=types.FunctionResponse(
                                id=part.function_call.id,
                                name=part.function_call.name,
                                response={"count": result},
                            )
                        )
                    ],
                ),
            ],
            config=config,
        )

Gemini never executes your SQL. Your app validates, runs, and returns structured JSON in functionResponse.

Log model id, function name, latency, and validation failures—OpenTelemetry traces help compare Gemini vs fallback models.

Parallel and Sequential Tool Turns

Gemini may return multiple functionCall parts when a user request decomposes naturally ("compare Q1 and Q2 revenue").

Parallel execution pattern:

calls = [p.function_call for p in parts if p.function_call]
results = await asyncio.gather(*[execute_safe(c) for c in calls])
# one functionResponse per call, each carrying that call's id

A compositional turn is the other shape: the second call needs the first result. Do not fan those out. Run the first, send its functionResponse, then let the model emit the next functionCall.

Parallel mutating tools need idempotency keys and per-tenant locks. Read-only analytics can fan out. Writes usually serialize.

Multi-turn tool chains follow the same loop as Tool Calling: cap max turns, summarize oversized results before re-injection.

Grounding vs Custom Tools

Google Search grounding and Maps grounding solve retrieval without you hosting APIs—they are not replacements for internal tools.

CapabilityGroundingCustom function
Internal warehouseNoYes
Public web factsYes (Search)Overkill
Auth to SaaSNoYes via your executor
Schema validationN/ARequired

Use grounding for external facts and function declarations for systems of record.

When one request mixes a built-in tool (Search, code execution) with a custom function, return every part Gemini sent, including the encrypted thought signature, plus your functionResponse. Dropping the signature makes the next turn lose the built-in tool context. The combination flow is in Using tools with the Gemini API.

Architecture Sketch

[Web app] --> [Orchestrator]
                  |
                  v
            [Gemini API / Vertex]
                  |
         functionCall parts
                  v
            [Tool executor] --> [DB / SaaS / queues]
                  ^
                  |-- validate, secrets, timeout

Production stacks keep API keys off the browser, route through your backend, and never trust raw model JSON without schema checks.

Reliability practices from Google SRE apply: alert when function invocation rate drops after model upgrades.

Readiness Scorecard

Rate readiness (1 point each):

CheckPass?
Function declarations ≤12 per active session
functionResponse includes the call id
Mode is AUTO or ANY, not NONE, on the paths that must call
Server-side validation before any mutating execute
AI Studio vs Vertex decision documented
Model id pinned in deploy config
Parallel call policy defined (read vs write)
Secrets in Secret Manager—not client bundle
Contract tests: 10+ prompts → expected function names
Oversized tool results summarized before re-injection
Observability: invocation rate, p95 latency, validation errors

8–12: production Gemini agents. 5–7: pilot one workflow. Below 5: demo—fix the declaration and the response loop first.

Cross-check UK NCSC guidelines for secure AI system development when Gemini reads customer data.

Failure Modes

Failure 1: The call never fires

The model returns prose and no functionCall. Check four things before you change models: the declaration is inside tools, the mode is not NONE, the parameter schema matches the question, and the previous turn sent functionResponse with the same id.

Failure 2: Missing functionResponse turn

Model repeats questions or hallucinates numbers. Fix: always inject a structured response, with the same id, before the next generate_content.

Failure 3: Bloated tool catalog

Wrong function selected. Fix: phase-based tool lists; improve descriptions.

Failure 4: AI Studio keys in production VPC

Compliance blockers. Fix: migrate to Vertex IAM.

Failure 5: Treating grounding as warehouse access

Public facts leak into internal decisions. Fix: separate grounding config from business tools.

Failure 6: Unbounded parallel writes

Race conditions on shared records. Fix: serialize mutating tools; idempotency keys.

Failure 7: Schema drift

Gemini sends new argument keys after an API change. Fix: strict validation and a structured error the model can read.

Operating Model

The integration needs one owner:

  • Maintain declaration JSON in git beside model config
  • Weekly review: function selection accuracy, validation error rate, p95 latency
  • Re-run contract tests on Gemini model id changes
  • Document rollback model id and declaration version
WeekFocus
1Two read-only tools + response-turn tests
2Executor validation + logging
3Vertex migration if required + IAM
4Parallel read tools + load test

Fifteen minutes weekly on mis-selected function names catches declaration drift before users notice.

InfiniSynapse Connection

InfiniSynapse optional layer for data-heavy tools behind Gemini: route long warehouse jobs to InfiniSynapse Server API while Gemini handles conversational tool selection. Your orchestrator keeps declarations; InfiniSynapse owns async compute and artifact download.

See LLM Tool Calling for multi-vendor routing and Agent Workflow Memory for summarizing large tool results.

Case Study: Workspace Admin Bot

A B2B SaaS team shipped an admin copilot on Gemini 2.0 Flash after a failed OpenAI-only prototype.

Path: six function declarations (list_tenants, usage_summary, open_ticket, etc.), AI Studio for week-one demos, Vertex migration with a service account for production. Added 14 fixture prompts asserting correct function names.

Results after five weeks:

  • Function selection accuracy: 78% → 93% after splitting read/write catalogs
  • Median read-tool latency: 1.4s → 620ms (parallel usage_summary + list_tenants)
  • Validation-blocked unsafe calls: 41/week → 3/week after enum constraints
  • Vertex migration: 2 engineer-days; zero declaration JSON changes
  • Incident: one missing functionResponse caused duplicate billing queries—fixed with integration test

The missing-response bug is why turn-by-turn tests beat prompt tweaking.

Post-launch dashboard: function name histogram, validation failures, Gemini vs fallback model usage. Selection accuracy is the number to compare with your OpenAI baseline.

Streaming and partial function calls

Gemini streaming can emit partial functionCall args before the turn completes. Buffer until the SDK marks the part finished. Never execute on half-formed JSON. Log part.thought separately from tool parts when using thinking models so support can tell "still reasoning" from "stuck without tools."

Multi-modal tool inputs

When users attach PDFs or screenshots, Gemini may call extraction tools before analytics tools. Treat vision inputs as state: store blob ids in orchestrator memory, pass only summaries into follow-up turns. This keeps token use predictable and mirrors patterns in Agent Workflow Memory.

Evaluation checklist before launch

Run twenty held-out prompts spanning paraphrases, typos, and adversarial "ignore tools" injections. Track precision/recall on expected function names—not just final natural-language answers. Teams that skip this step rediscover the same misses in production within the first sprint.

Frequently Asked Questions

What is a gemini tool call?

JSON the model asks you to run. You declared the function. Gemini returns functionCall. Your app executes it and sends functionResponse with the same id.

Does Gemini use the same schema as OpenAI tools?

Not exactly. Gemini uses function_declarations and functionResponse parts. The wire format differs. See Google AI function calling.

AI Studio or Vertex for production?

Prototype in AI Studio. Move to Vertex when you need IAM, regional pinning, or Cloud Logging. Do that before the tool can see customer data.

Can Gemini call multiple tools at once?

Yes—handle parallel functionCall parts with safe fan-out. Read-only analytics parallelize; serialize writes.

Does Gemini execute my Python functions?

No—it returns calls; your backend validates and executes, same as Tool Calling.

First step this week?

Declare one read-only function, run five prompts, confirm functionCall + your functionResponse loop end-to-end.

How does this compare to Claude or GPT tool use?

See LLM Tool Calling for the cross-vendor comparison. Gemini's practical strengths are Flash latency and a GCP-native path.

Conclusion

The loop is declaration discipline plus a functionResponse that carries the same id. Use Vertex when you need IAM. Validate arguments on the server every time.

Priority order: tighten declarations, test functionResponse turns, pin model ids, add observability, then expand parallel read tools.

Ship Gemini tools with fixture tests—not hope the model guesses your schema from vague descriptions.

Gemini Tool Call: Return the Function, Then Run It