Table of Contents
- The Problem with Feature Lists
- How Multimodal Input Actually Works
- The Think-Act-Observe Loop
- Skills vs. Examples: The 21% to 95% Jump
- Memory and State Management
- The Complete Pipeline
- What This Means for Your Architecture
The Problem with Feature Lists
Most coverage of Anthropic's capabilities reads like a product brochure. Claude can process images. It has tool use. It supports multi-turn conversations. All true, but completely useless if you're trying to build a production system.
The real question is: how do these pieces actually connect? When you feed an image and text prompt into Claude, invoke a database tool, maintain conversation context across twelve turns, and coordinate three specialized agents, what's actually happening at each boundary?
That's the map we're drawing here. Not what each component does in isolation, but how they wire together in a functioning system.
How Multimodal Input Actually Works
The term "multimodal" sounds impressive until you realize the model doesn't actually see pixels the way you do. Here's what really happens.
When you send an image to Claude, the vision encoder doesn't process every pixel. Instead, it learns a set of queries (typically 128 or 256) that act as attention probes across the image. Think of these as questions the model asks itself: "What's the most important information in this visual input?"
These queries scan the full pixel representation and compress it into a compact feature set. The core language model then processes these features alongside your text tokens. The image becomes just another sequence in the input stream.
Analogy: It's like asking someone to summarize a 200-page document into ten bullet points. The original information exists, but you're working with a distilled representation that captures what matters.
This compression matters for architects because it means:
- Image inputs consume a predictable token budget
- You can mix multiple images with text without exponential cost growth
- The model reasons over semantic features, not raw pixels
For a production pipeline, this means you can treat multimodal inputs as a preprocessing step that normalizes everything into the token space where your agent logic operates.
The Think-Act-Observe Loop
The core pattern in Anthropic's agentic systems is deceptively simple: think, act, observe. Each turn follows this sequence:
- Think: The model generates a brief reasoning trace
- Act: Issues exactly one tool call or returns the final answer
- Observe: Receives tool output as the next turn's observation
The critical constraint is one action per turn. This isn't a limitation, it's a design choice that makes the system traceable. Each decision point is individually attributable. When something goes wrong twelve steps into a multi-agent workflow, you can pinpoint exactly which reasoning step produced which action.
This matters more than it sounds. In parallel agent systems where multiple models can act simultaneously, debugging becomes archaeology. The one-action-per-turn design gives you a linear trace through even complex workflows.
| Turn | Think | Act | Observe |
|---|---|---|---|
| 1 | "Need customer data" | query_database(customer_id) | {name, history, preferences} |
| 2 | "Check inventory" | check_stock(product_id) | {available: true, quantity: 15} |
| 3 | "Calculate discount" | apply_rules(customer_tier) | {discount: 0.15} |
| 4 | "Ready to respond" | final_answer() | [end] |
Notice how each observation feeds directly into the next thinking step. State flows forward explicitly, not through hidden context.
Skills vs. Examples: The 21% to 95% Jump
Here's where most implementations fail. Anthropic's own testing found that agents without skills never exceeded 21% accuracy. With skills, the same systems consistently hit 95%+.
What's a skill? Not an example query. Not a few-shot prompt. A skill is a structured capability definition that tells the model:
- What this tool does (semantic description)
- When to use it (decision criteria)
- How to interpret results (output schema)
- What errors mean (failure modes)
The difference between examples and skills is the difference between "here's how someone used this once" and "here's the decision logic for when and how to use this."
For a database query tool, an example shows a sample query. A skill defines:
Name: customer_lookup
Use when: Need customer information for personalization or history
Input: customer_id (string, required)
Output: {name, tier, purchase_history[], preferences{}}
Errors: 404 = customer not found, 403 = access denied
Next steps: Use tier for pricing, history for recommendations
That additional context is what drives the accuracy jump. The model doesn't have to infer intent from examples; it has explicit decision criteria.
Memory and State Management
Multi-turn agentic systems need three types of memory:
Conversation memory: The full message history. This is your context window, managed by the API's message array. Every think-act-observe cycle appends to this log.
Working memory: Intermediate state within a single task. Results from tool calls that inform subsequent decisions. This lives in the observation field of each turn.
Long-term memory: Facts that persist across conversations. Customer preferences, learned patterns, system configuration. This requires external storage with retrieval tools.
The mistake is trying to cram everything into conversation memory. Context windows are large but not infinite. A 200-turn conversation with tool outputs will hit limits.
The right pattern:
- Conversation memory: recent turns + current task context
- Working memory: active tool results
- Long-term memory: retrieved on demand via tools
You summarize or drop old turns. You clear working memory after task completion. You store learned facts externally and retrieve them when relevant.
The Complete Pipeline
Here's how data flows through a real system:
Input Layer: Images and text arrive. Vision encoders compress images to feature queries. Text tokenizes normally. Everything becomes a token sequence.
Model Layer: Claude processes the unified token stream. Skills provide decision logic. The model generates reasoning traces and selects tools.
Tool Layer: Each turn executes one action. Tools return observations. Results feed back as context for the next turn.
Memory Layer: Conversation history grows. Working memory holds active results. Long-term memory retrieves via tools when needed.
Orchestration: Multiple specialized agents coordinate. Each follows the same think-act-observe pattern. A controller routes between agents based on task requirements.
The key insight: each layer has a clear contract. Multimodal inputs normalize to tokens. The model operates on tokens and skills. Tools consume actions and return observations. Memory provides context. Orchestration routes between independent agent loops.
What This Means for Your Architecture
If you're designing a production agentic system on Anthropic's stack, here's what matters:
Design for traceability first. The one-action-per-turn constraint isn't a bug. It gives you a linear audit trail. Embrace it.
Invest in skills, not examples. The 21% to 95% accuracy jump is real. Write proper skill definitions with decision criteria, not sample queries.
Separate memory types. Don't dump everything into conversation context. Use the right storage for each type: conversation, working, or long-term.
Normalize inputs early. Whether you start with images, PDFs, or structured data, get everything into token space before your agent logic runs.
Keep agents focused. Each agent should have a specific skill set and decision domain. Coordinate through a controller, don't try to build one mega-agent.
The stack isn't complicated. Multimodal encoding, model inference, tool execution, memory management, agent coordination. Five layers with clean contracts between them.
What makes it powerful is how those contracts compose. Once you map the boundaries, you can build systems that actually work in production instead of demos that impress in slides but fail in practice.
That's the real value of thinking in systems: not what each piece does, but how the whole thing flows.