Introducing the Model Context Protocol (MCP) The Model Context Protocol (MCP) is a JSON‑RPC based contract that lets a ChatGPT desktop agent discover and invoke “tools” running on a local or remote server. It bridges the LLM’s generic tool‑calling output and the concrete implementation of proprietary APIs, while keeping the LLM sandboxed from direct network access.
How MCP enables ChatGPT desktop agents to use proprietary tools
MCP defines a universal interface that a desktop client can query for available tools, send a JSON‑RPC request containing the tool name and arguments, and receive a structured response. The client (the MCP client library) translates a natural‑language instruction from the user into a tool call, forwards it to the MCP server, and injects the server’s response back into the LLM context. This flow provides a secure, auditable bridge between the LLM and internal systems without exposing those systems directly to the internet.
Prerequisites
- Python 3.11 or newer
- `mcp-sdk` (pip install mcp-sdk) – the official Python SDK for building MCP servers
- Access token for the internal REST API you intend to wrap (e.g., OAuth 2.0 client credentials)
- ChatGPT desktop app version 2.5 or newer (the built‑in MCP client is enabled by default)
Core architecture & concepts
MCP consists of three logical pieces: the desktop agent (client), the transport layer, and the MCP server (provider). The desktop agent runs inside the ChatGPT desktop process and uses the `mcp-client` library to locate server definitions in a JSON configuration file. Transport can be a simple stdio pipe, a TCP socket, or Server‑Sent Events (SSE) for remote servers. The server implements the JSON‑RPC methods `listTools`, `runTool`, and optional `getResource`.
graph LR
A[ChatGPT desktop] -->|mcp-client| B[Transport (stdio / SSE / TCP)]
B --> C[MCP server process]
C --> D[Tool implementation (Python, Java, etc.)]
D --> E[Internal API / system resource]
**Desktop agent as MCP client** – The client translates the LLM’s “call tool X with arguments Y” into a JSON‑RPC payload, sends it over the chosen transport, and waits for a response. The client also maintains a per‑chat session cache that stores tool results flagged as `cacheable`. This cache prevents accidental data leakage between unrelated chat sessions.
**Platform‑specific computer‑use APIs** – On Windows, macOS, and Linux the desktop agent can launch external programs, capture screenshots, or read the clipboard as defined by the MCP spec. Those capabilities are exposed to the LLM as separate “computer‑use” tools, but they are still invoked through the same JSON‑RPC channel.
**State flow** – When the LLM asks for a tool, the client checks the session cache. If a previous call in the same chat produced a result that is still valid, the cached value is returned without contacting the server. The cache key combines the tool name, argument hash, and the current user’s authentication scope, ensuring isolation between users and sessions.
Building and implementing a custom MCP server
1. Choose the development style
The MCP SDK provides a decorator‑based API that abstracts raw JSON‑RPC handling. If you prefer complete control you can implement the protocol by reading the specification and handling HTTP requests yourself, but the SDK reduces boilerplate and automatically validates request schemas.
# server.py – minimal MCP server using the SDK
from mcp_sdk import MCPServer, tool, jsonrpc
app = MCPServer(name="HR‑tool‑server", version="1.0.0")
@tool(name="get_employee", description="Fetch employee record by ID")
def get_employee(employee_id: str) -> dict:
"""Call the internal HRMS API and return JSON."""
return hrms_api.get(f"/employees/{employee_id}")
if __name__ == "__main__":
# stdio transport is the default for local desktop agents
app.run()
The code above is fully functional; running `python server.py` starts a process that reads JSON‑RPC messages from stdin and writes responses to stdout. The `@tool` decorator registers the function with the MCP catalog.
2. Wrap a private, authenticated REST API
Assume the HRMS API uses OAuth 2.0 client‑credentials flow with short‑lived access tokens. The SDK does not manage token refresh, so we implement a tiny helper that lazily fetches a token and retries on 401.
# hrms_client.py
import time
import requests
from typing import Dict
TOKEN_URL = "https://auth.example.com/oauth2/token"
API_BASE = "https://hrms.example.com/api"
CLIENT_ID = "my-client-id"
CLIENT_SECRET = "my-secret"
TOKEN_TTL = 300 # seconds
_token_cache: Dict[str, float] = {"access_token": "", "expires_at": 0.0}
def _fetch_token() -> str:
resp = requests.post(
TOKEN_URL,
data={"grant_type": "client_credentials"},
auth=(CLIENT_ID, CLIENT_SECRET),
timeout=5,
)
resp.raise_for_status()
data = resp.json()
_token_cache["access_token"] = data["access_token"]
_token_cache["expires_at"] = time.time() + data.get("expires_in", TOKEN_TTL)
return _token_cache["access_token"]
def _get_valid_token() -> str:
if time.time() >= _token_cache["expires_at"]:
return _fetch_token()
return _token_cache["access_token"]
def get(path: str) -> dict:
"""Perform a GET request with automatic token refresh."""
url = f"{API_BASE}{path}"
headers = {"Authorization": f"Bearer {_get_valid_token()}"}
resp = requests.get(url, headers=headers, timeout=5)
if resp.status_code == 401:
# token likely expired – refresh once and retry
headers["Authorization"] = f"Bearer {_fetch_token()}"
resp = requests.get(url, headers=headers, timeout=5)
resp.raise_for_status()
return resp.json()
Now integrate the client into the MCP server:
# server.py (continued)
from hrms_client import get as hrms_get
@tool(name="search_employees", description="Search employees by name fragment")
def search_employees(query: str, limit: int = 5) -> list:
"""Calls /employees/search?q=… and returns a list of matches."""
payload = hrms_get(f"/employees/search?q={query}&limit={limit}")
return payload.get("results", [])
The server now presents two tools (`get_employee` and `search_employees`) that the desktop agent can invoke. All token handling and error conversion happen inside the `hrms_client` module, keeping the tool functions pure.
3. Safety guardrails
MCP servers must validate incoming arguments because the LLM can synthesize malformed data. The SDK uses type hints to generate JSON schemas, but you can add explicit checks:
@tool(name="update_salary", description="Update salary for an employee")
def update_salary(employee_id: str, amount: float) -> str:
if amount <= 0:
raise jsonrpc.JSONRPCError(
code=-32602, message="amount must be positive"
)
resp = requests.post(
f"{API_BASE}/employees/{employee_id}/salary",
json={"amount": amount},
headers={"Authorization": f"Bearer {_get_valid_token()}"},
timeout=5,
)
resp.raise_for_status()
return "Salary updated"
Raising `jsonrpc.JSONRPCError` causes the MCP client to surface a clear error message to the LLM, which can then ask the user for clarification.
Connecting your MCP server to a desktop agent
Transport selection
- **Local stdio** – Fastest latency, suitable when the server runs on the same workstation.
- **SSE over HTTP** – Allows the server to run on a remote host while preserving a streaming response channel.
- **TCP socket** – Useful for containerised deployments behind a firewall.
The configuration file (`desktop-mcp-servers.json`) tells the ChatGPT desktop where to find the server. Example for a local stdio server:
{
"servers": [
{
"name": "HR‑tool‑server",
"command": "python /opt/mcp/servers/server.py",
"transport": "stdio",
"environment": {
"CLIENT_ID": "my-client-id",
"CLIENT_SECRET": "my-secret"
}
}
]
}
Place the file under `~/.config/ChatGPT/desktop-mcp-servers.json`. The desktop agent reads it at startup and automatically registers the server.
Verifying the integration
The MCP Inspector (bundled with the desktop app) lists all discovered tools and lets you invoke them manually. Open the inspector, select “HR‑tool‑server”, and run `search_employees` with a test query. The response should appear in the inspector UI and also be returned to the LLM if you ask:
User: Find employees whose name contains "Smith".
ChatGPT: (calls search_employees) → "Alice Smith, Bob Smith"
If the inspector shows a “malformed JSON‑RPC response” error, inspect the server logs for uncaught exceptions. The SDK prints a stack trace to stderr, which the desktop agent captures and reports.
Practical use cases and examples
Example 1 – HR agent with proprietary HRMS APIs
An internal help‑desk bot can retrieve employee profiles, start onboarding workflows, or reset passwords. The MCP server wraps the HRMS’s REST endpoints, adds token refresh logic, and enforces role‑based access by checking the `user_id` that the desktop client passes in the `metadata` field of every request.
Example 2 – Finance agent connecting to internal data sources
A finance analyst asks “What was the revenue for Q2 last year?”. The MCP server queries a private Snowflake warehouse via a Python connector, formats the result as a JSON table, and returns it. Because the query runs inside the secure corporate network, no credentials ever leave the perimeter.
Design patterns – multi‑agent chaining via resources
When a workflow requires several tools (e.g., fetch employee → calculate tenure → send Slack notification), the LLM can chain calls. The MCP spec encourages sharing intermediate data through “resources”. A server can expose a `get_resource` method that stores temporary blobs keyed by a UUID. Subsequent tool calls can retrieve the blob, allowing safe reuse without re‑executing expensive operations.
Key limitations, security, and best practices
Rate limits, timeouts, and partial failures
- Set a per‑tool timeout of 10 seconds in the server code (`requests.timeout`).
- Return a partial result with a `warning` field if the downstream API reports throttling; the client will surface the warning to the LLM.
- Implement exponential back‑off when retrying token refresh to avoid hammering the auth server.
Scoped credentials and sandboxed execution
Never run the MCP server as root. Use a dedicated system user with only the permissions required for the target API. Pass secrets via environment variables; the desktop client does not log them. The SDK sanitises any exception messages before they reach the LLM, preventing accidental leakage of stack traces.
Performance – latency, streaming, and concurrency
- **Latency** – Stdio transport adds ≈ 10 ms of overhead; SSE adds ≈ 30 ms due to network round‑trip. For sub‑second responsiveness, keep the server on the same host.
- **Streaming** – If a tool produces large output (e.g., log tail), implement the `runTool` method as a generator and return SSE chunks. The desktop client will forward each chunk to the LLM as it arrives.
- **Concurrency** – The SDK spawns a new thread per incoming request. Ensure all shared resources (e.g., token cache) are thread‑safe.
Trade‑offs & when not to use this approach
| Situation | Use MCP | Avoid MCP |
|---|---|---|
| Need to expose an internal API to many desktop users | ✅ Standardised interface, audit logs | ❌ Adds extra process hop, introduces latency |
| One‑off script that runs locally | ❌ Direct HTTP call is simpler | ✅ No additional server maintenance |
| Requirement for strict sandboxing and credential isolation | ✅ Server runs under limited OS user, credentials never leave host | ❌ Might be overkill for trusted internal scripts |
**Decision rule** – Choose MCP when you need a reusable, centrally managed bridge that enforces security policies. Choose direct API calls when the cost of an extra process outweighs the benefits.
Common errors and fixes
Error: “Failed to start MCP server process: No such file or directory”
**Cause** – The `command` path in `desktop-mcp-servers.json` is incorrect or the Python executable is not on the system PATH.
**Fix** – Use an absolute path to the interpreter or create a wrapper script.
{
"command": "/usr/bin/python3 /opt/mcp/servers/server.py"
}
Error: “JSON‑RPC response is not a valid object”
**Cause** – The tool function raised an exception that was not caught, causing the SDK to write a raw traceback to stdout.
**Fix** – Wrap the body in `try/except` and raise `jsonrpc.JSONRPCError` for known validation problems.
@tool(name="delete_employee")
def delete_employee(employee_id: str) -> str:
try:
hrms_get(f"/employees/{employee_id}/delete")
return "Deleted"
except requests.HTTPError as e:
raise jsonrpc.JSONRPCError(code=-32000, message=str(e))
Error: “Insufficient permission for tool get_employee”
**Cause** – The desktop client did not forward the user’s authentication scope, or the server ignored the `metadata` field.
**Fix** – Enable the `pass_metadata=True` flag when constructing the `MCPServer`, and inspect `metadata[“user_id”]` inside each tool to enforce ACLs.
app = MCPServer(name="HR-tool-server", version="1.0.0", pass_metadata=True)
Frequently asked questions
Can I connect an existing proprietary Python tool directly to ChatGPT without rewriting it as an API?
Yes. The MCP Python SDK lets you expose any callable as a tool. Register the function with the `@tool` decorator and the SDK handles the JSON‑RPC boundary, so no separate HTTP service is required.
What is the performance overhead of adding an MCP server layer versus a direct API integration?
Typical overhead is under 50 ms for stdio transport and under 100 ms for remote SSE. The dominant latency is still the underlying tool execution, not the MCP layer.
How do I share my custom MCP server configuration with my team for a desktop agent?
Commit the JSON file (`desktop-mcp-servers.json`) to your repo and distribute a short setup script that copies it to `~/.config/ChatGPT/`. The script can also set required environment variables.
Can I use MCP with a remote server behind a firewall?
Yes. Deploy the server behind the firewall, expose an HTTP endpoint that streams Server‑Sent Events, and set `transport: “sse”` in the client configuration. Ensure TLS termination at the edge.
How does the desktop agent keep state between tool calls?
The agent stores a per‑session cache keyed by tool name, argument hash, and user scope. Cached entries are cleared when the chat ends or when the tool signals `cacheable: false` in its response.
Practical next steps
Deploy the MCP server on a staging machine, register it with the desktop client, and use the MCP Inspector to validate each tool. Once the flow works end‑to‑end, add role‑based checks, enable resource sharing for multi‑step workflows, and monitor latency with a simple Prometheus exporter. With those pieces in place, your organization can expose internal APIs to ChatGPT desktop agents safely and consistently.