Tool Calling & Multimodal
Modern models can do more than read and write text. They can call tools you give them, and they can take images and audio as input. Knowing when to reach for these is what unlocks real features.
The Model Can Call a Function
You describe functions the model may use. Faced with a request, the model can decide to call one, read back the result, and keep going.
- You define the tools — look something up, do math, hit an API
- The model chooses when a tool is needed, and with what inputs
- It continues with the returned result folded into its answer
- This is the basis of agents — the whole of Module 7
Act on Real Data
A model on its own can only draw on what it saw in training — so for anything fresh or specific, it guesses. Tool calling lets it reach out and fetch the real thing instead.
Models That See and Hear
Some models accept images, and some audio, as input alongside text — not just words.
- Screenshots — describe or debug what is on a screen
- Documents — read a scanned page or a form
- Diagrams — interpret a chart or a hand sketch
- Reach for it only when the task genuinely needs more than text
Match the Capability to the Task
A bigger, multimodal model costs more to run — every request. If the job is plain text, use a plain-text model. Capability you don’t use is cost you still pay.
Build It
How to implement: find one place in your build where a tool would let the model call a function — say, “look up this week’s tasks” — and sketch that tool’s inputs and outputs.
- Weekly AI Tasks tracker — tool calling lets the model add and query tasks directly; a photo of a whiteboard could even become tasks (multimodal).
- Personal brand site — a multimodal model could draft alt-text from your project screenshots.
What you learned
Tool calling lets a model act on real data by calling functions you define — the basis of agents. Multimodal models take images and audio in. Match the capability to the task: unused capability is wasted cost.