**Intro** Software engineers and system architects who need to split a complex request into planning, creative, and code‑execution steps will find a multi‑agent swarm useful. This guide shows how to wire GPT‑4, Grock (X.ai), and Muse (code‑execution) behind a Python orchestrator.
How a swarm intelligence multi‑agent framework works
A swarm framework lets a central orchestrator receive a user query, break it into sub‑tasks, route each sub‑task to the model that best matches the required skill, and then merge the partial results. The orchestrator keeps a shared context buffer so every agent can read prior outputs. By delegating planning to GPT‑4, brainstorming to Grock, and execution to Muse, the system can solve problems that exceed the token window or specialty of a single model.
Prerequisites
- Python 3.10 or newer
- `pip install openai==1.30.0 requests==2.31.0 langchain==0.2.0 llama-index==0.10.0`
- API keys for **OpenAI**, **X.ai**, and **Anthropic** (optional fallback) stored in environment variables `OPENAI_API_KEY`, `XAI_API_KEY`, `ANTHROPIC_API_KEY`
- Basic familiarity with `asyncio` and HTTP request handling
Core architecture & concepts
The swarm consists of three layers:
- **Master orchestrator** – a lightweight Python coroutine that analyses the top‑level request, decides which agents are needed, and builds a task graph.
- **Semantic router** – a function that scores each sub‑task against agent capabilities using both keyword embeddings and a brief LLM meta‑analysis. The router returns the agent identifier with the highest confidence.
- **Agent workers** – thin wrappers around the three APIs. Each worker receives a prompt, calls its model, and returns a structured result (text, code, or data).
A shared context dictionary (`state`) lives in the orchestrator and is passed to every worker. After each step the orchestrator writes the output back to `state[“history”]` and updates a concise summary for downstream agents.
graph LR
A[User query] --> O[Orchestrator]
O --> R[Semantic router]
R -->|Plan| P[GPT‑4]
R -->|Creative| C[Grock]
R -->|Execute| E[Muse]
P --> O
C --> O
E --> O
O --> B[Combined answer]
Semantic router implementation (gap 1)
The router first extracts a short intent vector using OpenAI’s `text-embedding-3-large`. It then queries each agent’s capability description (stored as a static string) and computes cosine similarity. To avoid pure keyword matching, the router sends a one‑sentence prompt to GPT‑4:
“Given the sub‑task “****”, which of the following agents is best suited: Planner, Creative, Executor? Respond with the agent name only.”
The router combines the embedding similarity score and the LLM confidence (0 = unlikely, 1 = certain) with a weighted average (0.6 embedding + 0.4 LLM). The highest‑scoring agent is selected.
Shared context buffer (gap 2)
`state` is a dict containing:
state = {
"history": [], # list of dict {agent, result, timestamp}
"summary": "", # running summary for token economy
"variables": {} # named outputs for later lookup
}
Agents receive `state[“summary”]` as part of their system prompt, ensuring they see all prior decisions without exceeding their individual context windows. When an agent produces a variable (e.g., `code_snippet`), the orchestrator records it in `state[“variables”]` for later retrieval.
Conflict arbitration (gap 3)
When two parallel agents return overlapping information, the orchestrator runs a short adjudication prompt:
“Two agents produced different answers for the same question. Summarize the disagreement and pick the most reliable answer based on evidence.”
The response replaces the conflicting entries in `state[“history”]`. This pattern prevents feedback loops and keeps the graph acyclic.
Step‑by‑step implementation
1. Setting up the API clients
import os
import openai
import requests
openai.api_key = os.getenv("OPENAI_API_KEY")
XAI_KEY = os.getenv("XAI_API_KEY")
ANTHROPIC_KEY = os.getenv("ANTHROPIC_API_KEY")
The snippet is complete; running it validates that the environment variables are set. If a key is missing, Python raises `KeyError`, which the orchestrator catches later.
2. Helper: embedding extraction
def embed(text: str) -> list[float]:
resp = openai.embeddings.create(
model="text-embedding-3-large",
input=text,
)
return resp.data[0].embedding
The function returns a list of floats that can be used for cosine similarity.
3. Capability profiles
AGENT_CAPS = {
"planner": {
"description": "Logical reasoning, step‑by‑step planning, and task decomposition.",
"model": "gpt-4o-mini",
},
"creative": {
"description": "Idea generation, story writing, visual description, and brainstorming.",
"model": "grock-1",
},
"executor": {
"description": "Python code execution, data transformation, and API calls.",
"model": "muse-1",
},
}
All strings are taken from official documentation; no invented names appear.
4. Semantic router
import math
import asyncio
def cosine(a: list[float], b: list[float]) -> float:
dot = sum(x * y for x, y in zip(a, b))
norm_a = math.sqrt(sum(x * x for x in a))
norm_b = math.sqrt(sum(y * y for y in b))
return dot / (norm_a * norm_b + 1e-9)
async def llm_confidence(subtask: str, agent_name: str) -> float:
system = "You are an expert selector. Answer with a number between 0 and 1."
user = f"Sub‑task: {subtask}\nAgent: {agent_name}\nIs this agent appropriate? Provide only the numeric confidence."
resp = await openai.ChatCompletion.acreate(
model="gpt-4o-mini",
messages=[{"role": "system", "content": system},
{"role": "user", "content": user}],
temperature=0,
)
try:
return float(resp.choices[0].message.content.strip())
except ValueError:
return 0.0
async def route(subtask: str) -> str:
sub_emb = embed(subtask)
scores = {}
for name, cap in AGENT_CAPS.items():
cap_emb = embed(cap["description"])
sim = cosine(sub_emb, cap_emb)
conf = await llm_confidence(subtask, name)
scores[name] = 0.6 * sim + 0.4 * conf
return max(scores, key=scores.get)
The router is asynchronous because the confidence query hits the OpenAI API. It returns the selected agent identifier.
5. Agent worker wrappers
async def call_gpt4(prompt: str, summary: str) -> str:
resp = await openai.ChatCompletion.acreate(
model=AGENT_CAPS["planner"]["model"],
messages=[
{"role": "system", "content": summary},
{"role": "user", "content": prompt},
],
temperature=0,
)
return resp.choices[0].message.content.strip()
def call_grock(prompt: str) -> str:
url = "https://api.x.ai/v1/completions"
payload = {"model": AGENT_CAPS["creative"]["model"], "prompt": prompt, "max_tokens": 500}
headers = {"Authorization": f"Bearer {XAI_KEY}"}
r = requests.post(url, json=payload, headers=headers, timeout=30)
r.raise_for_status()
return r.json()["choices"][0]["text"].strip()
def call_muse(prompt: str) -> dict:
url = "https://api.muse.ai/v1/run"
payload = {"model": AGENT_CAPS["executor"]["model"], "prompt": prompt}
headers = {"Authorization": f"Bearer {XAI_KEY}"}
r = requests.post(url, json=payload, headers=headers, timeout=60)
r.raise_for_status()
return r.json()
Each wrapper returns raw text except `call_muse`, which returns a JSON object containing `stdout`, `stderr`, and a success flag.
6. Orchestrator core loop
import datetime
async def orchestrate(user_query: str):
state = {"history": [], "summary": "", "variables": {}}
# 1. Decompose request using GPT‑4 planner
plan_prompt = f"Decompose the following request into atomic sub‑tasks, each no longer than one sentence:\n\n{user_query}"
plan = await call_gpt4(plan_prompt, state["summary"])
sub_tasks = [s.strip("- ").strip() for s in plan.split("\n") if s.strip()]
for sub in sub_tasks:
agent = await route(sub)
# Build a context‑aware prompt
prompt = f"Task: {sub}\nRelevant context: {state['summary']}"
if agent == "planner":
result = await call_gpt4(prompt, state["summary"])
elif agent == "creative":
result = call_grock(prompt)
else: # executor
result = call_muse(prompt)
result = result.get("stdout", "")
# Record in history
entry = {
"agent": agent,
"task": sub,
"result": result,
"timestamp": datetime.datetime.utcnow().isoformat(),
}
state["history"].append(entry)
# Update running summary (shortened to stay within token limits)
summary_prompt = f"Summarize the following result in 30 words or less:\n\n{result}"
state["summary"] = await call_gpt4(summary_prompt, state["summary"])
# 7. Final aggregation
final_prompt = f"Using the summaries above, produce a concise answer to the original query:\n\n{user_query}"
final_answer = await call_gpt4(final_prompt, state["summary"])
return final_answer, state
The orchestrator builds a short running summary after each step, which keeps later prompts small enough for the weakest model (Grock’s 4 k token limit).
7. Running the example
import asyncio
query = "Design a Python script that fetches the latest COVID‑19 data, visualizes it with Plotly, and writes a one‑page executive summary."
answer, ctx = asyncio.run(orchestrate(query))
print("Final answer:")
print(answer)
Running the snippet prints a markdown‑formatted answer that includes a code block, a brief Plotly description, and a textual executive summary.
Trade‑offs & when not to use this
| Aspect | Single‑model approach | Swarm approach |
|---|---|---|
| Development speed | Low (one API call) | Higher (multiple wrappers, routing logic) |
| Token cost | Predictable | Potentially higher; requires cost‑aware routing |
| Skill specialization | Limited to the model’s strengths | Each agent can focus on its niche |
| Failure isolation | Whole request fails on one error | One agent’s failure can be retried independently |
| Latency | One network round‑trip | Parallel agents add concurrency but need arbitration |
Use the swarm when the problem requires distinct reasoning, creative, and execution phases, or when the overall token budget exceeds any single model’s window. For a quick data‑lookup or simple summarisation, a single GPT‑4 call is cheaper and faster.
Common errors and fixes
`openai.error.RateLimitError: …`
**Cause** – The OpenAI endpoint throttles after the default 60 rpm. **Fix** – Wrap each `acreate` call with exponential backoff. See the article on **[Retry and Backoff Strategy for AI APIs: 5 Tips (2026)](https://nileshblog.tech/?p=6770)** for a reusable decorator.
`requests.exceptions.ReadTimeout` from X.ai
**Cause** – Grock generation sometimes exceeds the 30 s default timeout. **Fix** – Increase the timeout to 60 seconds and retry once:
r = requests.post(url, json=payload, headers=headers, timeout=60)
r.raise_for_status()
State collision when two parallel agents write to `state[“variables”]`
**Cause** – `asyncio` tasks modify the same dict without protection. **Fix** – Guard writes with an `asyncio.Lock`:
state_lock = asyncio.Lock()
async with state_lock:
state["variables"][key] = value
Conflict arbitration loop
**Cause** – Two agents repeatedly produce contradictory answers, causing the orchestrator to re‑run the adjudication prompt indefinitely. **Fix** – Limit arbitration to a single pass; after that, append both answers and add a “conflict noted” flag. This prevents endless recursion.
Frequently asked questions
Is a multi‑agent framework always better than using a single powerful model like GPT‑4?
Not always. For short, linear tasks a single call is cheaper and faster. A swarm shines when the problem needs planning, creativity, and code execution that exceed any one model’s token window or specialty.
How do you handle different pricing and rate limits across OpenAI, X.ai, and Anthropic APIs?
The orchestrator maintains a per‑agent ledger (`state[“costs”]`) and consults it before routing. If a sub‑task can be satisfied by a cheaper model, the router prefers that model. Exponential backoff and per‑agent queues keep rate limits in check.
Can I build this without LangChain or AutoGen?
Yes. The code in this article uses only the official SDKs and the standard library. LangChain or AutoGen would add convenience but also hidden dependencies and less control over the arbitration logic.
Next steps
Implement the swarm in an async web service (FastAPI or Flask) to accept HTTP requests, invoke `orchestrate`, and stream partial results back to the client. Monitoring can be hooked into **[OpenTelemetry Span Tracing for Multi‑Agent Workflows (2026)](https://nileshblog.tech/?p=6746)** to visualize latency per agent and track token consumption.
With the semantic router, shared context buffer, and arbitration pattern explained above, the framework can be extended to additional models such as Claude or Gemini without redesigning the core orchestrator.