**Intro** Software engineers who frequently test multiple large‑language‑model (LLM) providers need a quick way to change the target model without restarting a script. A small desktop dashboard built with Python keeps the workflow local, avoids a web server, and lets you compare responses side‑by‑side.
How a live model‑switching UI works
A live model‑switching UI is a Python desktop program that lets a user type a prompt, pick a provider (for example OpenAI‑GPT‑4 or Anthropic‑Claude‑Opus) from a dropdown, and see the answer appear while the request is still streaming. The UI stays responsive because the actual HTTP call runs in a background thread; the main thread only updates widgets via `after`. The router maintains a separate conversation history for each provider so that a switch does not lose context.
Prerequisites
| Item | Minimum version |
|---|---|
| Python | 3.10 |
| `openai` package | 1.13.0 |
| `anthropic` package | 0.6.0 |
| `tkinter` (bundled with CPython) | — |
| `requests` (used by SDKs) | 2.31.0 |
| `pyinstaller` (optional, for distribution) | 6.2.0 |
You also need valid API keys for OpenAI and Anthropic. Export them in the shell before running the script:
export OPENAI_API_KEY=sk-...
export ANTHROPIC_API_KEY=claude-...
Core architecture & concepts
The dashboard revolves around three logical layers:
- **Provider adapters** – thin wrappers that expose a unified `generate(prompt, history, params)` method.
- **ModelRouter** – stores per‑provider conversation history, selects the active adapter, and forwards the request.
- **Tkinter UI** – collects the user prompt, displays the streaming output, and triggers the router in a worker thread.
The data flow is illustrated with a Mermaid diagram:
flowchart LR
UI[Tkinter UI] -->|user prompt| Router[ModelRouter]
Router -->|selected adapter| OpenAI[OpenAIAdapter]
Router -->|selected adapter| Anthropic[AnthropicAdapter]
OpenAI -->|stream tokens| UI
Anthropic -->|stream tokens| UI
Router -->|update history| Store[Conversation Store]
Defining a standard interface
All adapters inherit from an abstract base class that defines the contract:
from abc import ABC, abstractmethod
from typing import List, Dict, Generator
class BaseAdapter(ABC):
@abstractmethod
def generate(self,
prompt: str,
history: List[Dict[str, str]],
params: Dict) -> Generator[str, None, None]:
"""Yield token chunks as they arrive."""
pass
@abstractmethod
def format_history(self,
history: List[Dict[str, str]]) -> List[Dict]:
"""Return provider‑specific message structure."""
pass
Implementing OpenAI and Anthropic modules
import openai
from anthropic import Anthropic, MESSAGE_DELIMITER
class OpenAIAdapter(BaseAdapter):
def __init__(self, model: str = "gpt-4o"):
self.model = model
def format_history(self, history):
return [{"role": h["role"], "content": h["content"]} for h in history]
def generate(self, prompt, history, params):
messages = self.format_history(history) + [{"role": "user", "content": prompt}]
response = openai.ChatCompletion.create(
model=self.model,
messages=messages,
temperature=params.get("temperature", 0.7),
max_tokens=params.get("max_tokens", 1024),
stream=True,
)
for chunk in response:
delta = chunk["choices"][0]["delta"]
if "content" in delta:
yield delta["content"]
class AnthropicAdapter(BaseAdapter):
def __init__(self, model: str = "claude-3-opus-20240229"):
self.client = Anthropic()
self.model = model
def format_history(self, history):
# Anthropic expects a list of {"role": "assistant/user", "content": "..."}
return [{"role": h["role"], "content": h["content"]} for h in history]
def generate(self, prompt, history, params):
messages = self.format_history(history) + [{"role": "user", "content": prompt}]
with self.client.messages.stream(
model=self.model,
max_tokens=params.get("max_tokens", 1024),
temperature=params.get("temperature", 0.7),
messages=messages,
) as stream:
for event in stream:
if event.type == "content_block_delta":
yield event.delta["text"]
Both adapters return a generator that yields token fragments. The router treats them identically.
Managing API keys and session state securely
The adapters rely on environment variables accessed by the SDKs; no key is ever written to disk. The `ModelRouter` keeps a `dict` keyed by provider name that stores:
- `history` – list of message dicts for that provider
- `adapter` – instantiated wrapper
import os
from queue import Queue
class ModelRouter:
def __init__(self):
self.providers = {
"openai": {
"adapter": OpenAIAdapter(),
"history": [],
},
"anthropic": {
"adapter": AnthropicAdapter(),
"history": [],
},
}
self.active = "openai"
self.response_queue = Queue()
def switch_provider(self, name: str):
if name not in self.providers:
raise ValueError(f"Unknown provider {name}")
self.active = name
def send_prompt(self, prompt: str, params: dict):
provider = self.providers[self.active]
generator = provider["adapter"].generate(
prompt, provider["history"], params
)
for token in generator:
self.response_queue.put(token)
# Append the completed interaction to history
provider["history"].append({"role": "user", "content": prompt})
provider["history"].append({"role": "assistant", "content": "".join(
list(self.response_queue.queue)
)})
# Clear the queue for the next request
while not self.response_queue.empty():
self.response_queue.get()
The router guarantees that each provider’s context remains independent, solving **Gap 2** from the brief.
Building the Tkinter dashboard UI
Prompt input and model selector
import tkinter as tk
from tkinter import ttk
class Dashboard(tk.Tk):
def __init__(self, router: ModelRouter):
super().__init__()
self.title("LLM Router")
self.geometry("720x540")
self.router = router
# Prompt entry
self.prompt = tk.Text(self, height=4, wrap="word")
self.prompt.pack(fill="x", padx=10, pady=5)
# Provider dropdown
self.provider_var = tk.StringVar(value="openai")
self.provider_menu = ttk.Combobox(
self,
textvariable=self.provider_var,
values=list(self.router.providers.keys()),
state="readonly",
)
self.provider_menu.pack(fill="x", padx=10, pady=5)
# Send button
self.send_btn = ttk.Button(self, text="Send", command=self.on_send)
self.send_btn.pack(pady=5)
# Output console
self.output = tk.Text(self, height=20, wrap="word", state="disabled")
self.output.pack(fill="both", expand=True, padx=10, pady=5)
def on_send(self):
# Switch router to selected provider
self.router.switch_provider(self.provider_var.get())
user_prompt = self.prompt.get("1.0", "end-1c")
self.prompt.delete("1.0", "end")
self.output.configure(state="normal")
self.output.insert("end", f"> {user_prompt}\n")
self.output.configure(state="disabled")
# Start background thread
threading.Thread(
target=self.run_generation,
args=(user_prompt,),
daemon=True,
).start()
Streaming token simulation
Tkinter cannot be updated from a non‑main thread. The background worker places each token into a thread‑safe queue, and the main loop polls the queue with `after`.
import threading
import queue
def run_generation(self, prompt):
params = {"temperature": 0.7, "max_tokens": 1024}
self.router.send_prompt(prompt, params)
self.after(10, self.poll_queue)
def poll_queue(self):
try:
token = self.router.response_queue.get_nowait()
except queue.Empty:
# No more tokens, re‑schedule poll
self.after(10, self.poll_queue)
return
self.output.configure(state="normal")
self.output.insert("end", token)
self.output.configure(state="disabled")
self.output.see("end")
# Continue polling until queue is empty
self.after(10, self.poll_queue)
The UI never freezes because the heavy network I/O happens in the `threading.Thread`. The `after` callback runs on the main thread, avoiding the `TclError: main thread is not in main loop` failure mode described later.
Implementing the event‑driven message loop
Threading the API calls
All API calls are wrapped in a daemon thread. Daemon threads exit automatically when the main program terminates, preventing orphan processes.
Managing pending requests and cancellation
A simple `threading.Event` can be used to abort a request if the user clicks **Cancel** (not shown in the minimal UI). The event is checked inside the generator loop; if set, the generator stops yielding.
class CancelableGenerator:
def __init__(self, gen):
self._gen = gen
self._cancel = threading.Event()
def cancel(self):
self._cancel.set()
def __iter__(self):
for item in self._gen:
if self._cancel.is_set():
break
yield item
Real‑time UI updates
The `poll_queue` method demonstrates a clean separation: the worker never touches Tkinter widgets, and the UI reads from a synchronized queue only.
Handling provider‑specific APIs and outputs
Standardizing JSON and function‑calling responses
Both OpenAI and Anthropic can return structured JSON when `response_format` (OpenAI) or `tool_choice` (Anthropic) is set. The router adds a `json_mode` flag to `params` and passes it unchanged; each adapter translates the flag to the provider’s exact syntax.
# OpenAI example
if params.get("json_mode"):
response = openai.ChatCompletion.create(
model=self.model,
messages=messages,
response_format={"type": "json_object"},
stream=True,
)
# Anthropic example
if params.get("json_mode"):
messages = self.format_history(history) + [
{"role": "user", "content": [{"type": "text", "text": prompt}]}
]
# Anthropic reads JSON from a tool call; simplified here
Parsing system prompts and conversation history
System prompts differ: OpenAI uses a `system` role, while Anthropic treats the first message as a system prompt if its role is `”assistant”` with a `”content”` field marked as `type: “system”`. The adapters implement `format_history` accordingly, preventing the **inconsistent conversation state** error when switching models.
Configuring per‑model parameters
The UI could expose sliders for `temperature` and `max_tokens`. The router forwards the exact dictionary; each adapter extracts the keys it supports, leaving unknown keys untouched.
Testing and debugging the live routing system
Simulating API errors and rate limiting
Wrap the generator in a try/except block inside `ModelRouter.send_prompt`. On `openai.RateLimitError` or `anthropic.RateLimitError`, put a descriptive token into the queue so the UI shows “Rate limit exceeded”.
try:
for token in generator:
self.response_queue.put(token)
except (openai.RateLimitError, anthropic.RateLimitError) as e:
self.response_queue.put(f"\n[Error] {e}")
Logging prompts and responses
A lightweight logger writes each interaction to a rotating file. This creates an audit trail required for production use.
import logging
logger = logging.getLogger("router")
handler = logging.handlers.RotatingFileHandler(
"router.log", maxBytes=1_048_576, backupCount=3
)
logger.setLevel(logging.INFO)
logger.addHandler(handler)
def log_interaction(provider, prompt, response):
logger.info(
f"{provider.upper()} | Prompt: {prompt!r} | Response: {response!r}"
)
The `log_interaction` call can be placed after the response queue is emptied.
Validating model outputs after a switch
After a provider switch, the UI can display a small banner summarizing the current conversation length for that provider. This lets the developer verify that the history count matches expectations.
def show_history_len(self):
hist = self.router.providers[self.router.active]["history"]
self.output.configure(state="normal")
self.output.insert("end", f"\n[Info] {self.router.active} history length: {len(hist)}\n")
self.output.configure(state="disabled")
Trade‑offs & when not to use this
| Aspect | Tkinter desktop router | Web‑based selector |
|---|---|---|
| Latency | Minimal, no network round‑trip for UI assets | Dependent on HTTP server and browser |
| Multi‑user support | Single user per process | Naturally multi‑user |
| Distribution | Packaged as an executable with PyInstaller | Requires hosting infrastructure |
| UI richness | Limited widgets, less styling | Full HTML/CSS/JS capabilities |
| Complexity | Simple thread‑based design | Needs async server, auth, scaling |
**When to choose the desktop router** – you need a quick, locally controlled tool for model comparison, you do not require simultaneous access by many users, and you prefer a single binary that can run on a developer workstation.
**When to avoid it** – you must share the interface across a team, need real‑time collaboration, or want to embed the router inside another web service.
Common errors and fixes
`TclError: main thread is not in main loop`
*Cause* – Attempting to modify a Tkinter widget from a background thread.
*Fix* – Always route UI updates through `after` or a thread‑safe queue. The `poll_queue` pattern above satisfies this rule.
Inconsistent conversation state after switching models
*Cause* – The new provider receives a history that contains roles it does not understand (e.g., OpenAI `system` messages sent to Claude).
*Fix* – Each adapter’s `format_history` must translate roles to the provider’s accepted schema. The code in the adapters already performs this conversion.
Rate‑limit exception crashes the GUI
*Cause* – Unhandled exception propagates out of the worker thread.
*Fix* – Catch provider‑specific rate‑limit errors inside `ModelRouter.send_prompt` and push an error token to the queue as shown earlier.
Frequently asked questions
How do I prevent the Tkinter GUI from freezing when calling an AI API?
Run the API call inside a `threading.Thread`. Return token fragments through a `queue.Queue` and let the main thread pull them with `after` for safe UI updates.
Can I compare GPT‑4 and Claude‑Opus outputs side‑by‑side?
Yes. Extend the UI with two read‑only text widgets and modify the router to dispatch the same prompt to both adapters simultaneously. Each worker thread writes to its own queue, which the corresponding output pane polls.
What is the best way to package this dashboard for non‑technical users?
Use `pyinstaller –onefile –windowed dashboard.py`. The generated executable bundles the Python interpreter, required wheels, and Tkinter runtime, allowing distribution without a separate Python installation.
How can I add a local LLM served by Ollama to the router?
Create a new adapter that posts to `http://localhost:11434/api/chat` using `requests`. Implement `generate` as a generator that yields `response[“message”][“content”]` chunks, then register the adapter in `ModelRouter.providers`.