**Intro** Software engineers who frequently test multiple large‑language‑model (LLM) providers need a quick way to change the target model without restarting a script. A small desktop dashboard built with Python keeps the workflow local, avoids a web server, and lets you compare responses side‑by‑side.

How a live model‑switching UI works

A live model‑switching UI is a Python desktop program that lets a user type a prompt, pick a provider (for example OpenAI‑GPT‑4 or Anthropic‑Claude‑Opus) from a dropdown, and see the answer appear while the request is still streaming. The UI stays responsive because the actual HTTP call runs in a background thread; the main thread only updates widgets via `after`. The router maintains a separate conversation history for each provider so that a switch does not lose context.

Prerequisites

ItemMinimum version
Python3.10
`openai` package1.13.0
`anthropic` package0.6.0
`tkinter` (bundled with CPython)—
`requests` (used by SDKs)2.31.0
`pyinstaller` (optional, for distribution)6.2.0

You also need valid API keys for OpenAI and Anthropic. Export them in the shell before running the script:

export OPENAI_API_KEY=sk-...
export ANTHROPIC_API_KEY=claude-...

Core architecture & concepts

The dashboard revolves around three logical layers:

  1. **Provider adapters** – thin wrappers that expose a unified `generate(prompt, history, params)` method.
  2. **ModelRouter** – stores per‑provider conversation history, selects the active adapter, and forwards the request.
  3. **Tkinter UI** – collects the user prompt, displays the streaming output, and triggers the router in a worker thread.

The data flow is illustrated with a Mermaid diagram:

flowchart LR
    UI[Tkinter UI] -->|user prompt| Router[ModelRouter]
    Router -->|selected adapter| OpenAI[OpenAIAdapter]
    Router -->|selected adapter| Anthropic[AnthropicAdapter]
    OpenAI -->|stream tokens| UI
    Anthropic -->|stream tokens| UI
    Router -->|update history| Store[Conversation Store]

Defining a standard interface

All adapters inherit from an abstract base class that defines the contract:

from abc import ABC, abstractmethod
from typing import List, Dict, Generator

class BaseAdapter(ABC):
    @abstractmethod
    def generate(self,
                 prompt: str,
                 history: List[Dict[str, str]],
                 params: Dict) -> Generator[str, None, None]:
        """Yield token chunks as they arrive."""
        pass

    @abstractmethod
    def format_history(self,
                       history: List[Dict[str, str]]) -> List[Dict]:
        """Return provider‑specific message structure."""
        pass

Implementing OpenAI and Anthropic modules

import openai
from anthropic import Anthropic, MESSAGE_DELIMITER

class OpenAIAdapter(BaseAdapter):
    def __init__(self, model: str = "gpt-4o"):
        self.model = model

    def format_history(self, history):
        return [{"role": h["role"], "content": h["content"]} for h in history]

    def generate(self, prompt, history, params):
        messages = self.format_history(history) + [{"role": "user", "content": prompt}]
        response = openai.ChatCompletion.create(
            model=self.model,
            messages=messages,
            temperature=params.get("temperature", 0.7),
            max_tokens=params.get("max_tokens", 1024),
            stream=True,
        )
        for chunk in response:
            delta = chunk["choices"][0]["delta"]
            if "content" in delta:
                yield delta["content"]

class AnthropicAdapter(BaseAdapter):
    def __init__(self, model: str = "claude-3-opus-20240229"):
        self.client = Anthropic()
        self.model = model

    def format_history(self, history):
        # Anthropic expects a list of {"role": "assistant/user", "content": "..."}
        return [{"role": h["role"], "content": h["content"]} for h in history]

    def generate(self, prompt, history, params):
        messages = self.format_history(history) + [{"role": "user", "content": prompt}]
        with self.client.messages.stream(
            model=self.model,
            max_tokens=params.get("max_tokens", 1024),
            temperature=params.get("temperature", 0.7),
            messages=messages,
        ) as stream:
            for event in stream:
                if event.type == "content_block_delta":
                    yield event.delta["text"]

Both adapters return a generator that yields token fragments. The router treats them identically.

Managing API keys and session state securely

The adapters rely on environment variables accessed by the SDKs; no key is ever written to disk. The `ModelRouter` keeps a `dict` keyed by provider name that stores:

  • `history` – list of message dicts for that provider
  • `adapter` – instantiated wrapper
import os
from queue import Queue

class ModelRouter:
    def __init__(self):
        self.providers = {
            "openai": {
                "adapter": OpenAIAdapter(),
                "history": [],
            },
            "anthropic": {
                "adapter": AnthropicAdapter(),
                "history": [],
            },
        }
        self.active = "openai"
        self.response_queue = Queue()

    def switch_provider(self, name: str):
        if name not in self.providers:
            raise ValueError(f"Unknown provider {name}")
        self.active = name

    def send_prompt(self, prompt: str, params: dict):
        provider = self.providers[self.active]
        generator = provider["adapter"].generate(
            prompt, provider["history"], params
        )
        for token in generator:
            self.response_queue.put(token)
        # Append the completed interaction to history
        provider["history"].append({"role": "user", "content": prompt})
        provider["history"].append({"role": "assistant", "content": "".join(
            list(self.response_queue.queue)
        )})
        # Clear the queue for the next request
        while not self.response_queue.empty():
            self.response_queue.get()

The router guarantees that each provider’s context remains independent, solving **Gap 2** from the brief.

Building the Tkinter dashboard UI

Prompt input and model selector

import tkinter as tk
from tkinter import ttk

class Dashboard(tk.Tk):
    def __init__(self, router: ModelRouter):
        super().__init__()
        self.title("LLM Router")
        self.geometry("720x540")
        self.router = router

        # Prompt entry
        self.prompt = tk.Text(self, height=4, wrap="word")
        self.prompt.pack(fill="x", padx=10, pady=5)

        # Provider dropdown
        self.provider_var = tk.StringVar(value="openai")
        self.provider_menu = ttk.Combobox(
            self,
            textvariable=self.provider_var,
            values=list(self.router.providers.keys()),
            state="readonly",
        )
        self.provider_menu.pack(fill="x", padx=10, pady=5)

        # Send button
        self.send_btn = ttk.Button(self, text="Send", command=self.on_send)
        self.send_btn.pack(pady=5)

        # Output console
        self.output = tk.Text(self, height=20, wrap="word", state="disabled")
        self.output.pack(fill="both", expand=True, padx=10, pady=5)

    def on_send(self):
        # Switch router to selected provider
        self.router.switch_provider(self.provider_var.get())
        user_prompt = self.prompt.get("1.0", "end-1c")
        self.prompt.delete("1.0", "end")
        self.output.configure(state="normal")
        self.output.insert("end", f"> {user_prompt}\n")
        self.output.configure(state="disabled")
        # Start background thread
        threading.Thread(
            target=self.run_generation,
            args=(user_prompt,),
            daemon=True,
        ).start()

Streaming token simulation

Tkinter cannot be updated from a non‑main thread. The background worker places each token into a thread‑safe queue, and the main loop polls the queue with `after`.

import threading
import queue

    def run_generation(self, prompt):
        params = {"temperature": 0.7, "max_tokens": 1024}
        self.router.send_prompt(prompt, params)
        self.after(10, self.poll_queue)

    def poll_queue(self):
        try:
            token = self.router.response_queue.get_nowait()
        except queue.Empty:
            # No more tokens, re‑schedule poll
            self.after(10, self.poll_queue)
            return
        self.output.configure(state="normal")
        self.output.insert("end", token)
        self.output.configure(state="disabled")
        self.output.see("end")
        # Continue polling until queue is empty
        self.after(10, self.poll_queue)

The UI never freezes because the heavy network I/O happens in the `threading.Thread`. The `after` callback runs on the main thread, avoiding the `TclError: main thread is not in main loop` failure mode described later.

Implementing the event‑driven message loop

Threading the API calls

All API calls are wrapped in a daemon thread. Daemon threads exit automatically when the main program terminates, preventing orphan processes.

Managing pending requests and cancellation

A simple `threading.Event` can be used to abort a request if the user clicks **Cancel** (not shown in the minimal UI). The event is checked inside the generator loop; if set, the generator stops yielding.

class CancelableGenerator:
    def __init__(self, gen):
        self._gen = gen
        self._cancel = threading.Event()

    def cancel(self):
        self._cancel.set()

    def __iter__(self):
        for item in self._gen:
            if self._cancel.is_set():
                break
            yield item

Real‑time UI updates

The `poll_queue` method demonstrates a clean separation: the worker never touches Tkinter widgets, and the UI reads from a synchronized queue only.

Handling provider‑specific APIs and outputs

Standardizing JSON and function‑calling responses

Both OpenAI and Anthropic can return structured JSON when `response_format` (OpenAI) or `tool_choice` (Anthropic) is set. The router adds a `json_mode` flag to `params` and passes it unchanged; each adapter translates the flag to the provider’s exact syntax.

# OpenAI example
if params.get("json_mode"):
    response = openai.ChatCompletion.create(
        model=self.model,
        messages=messages,
        response_format={"type": "json_object"},
        stream=True,
    )
# Anthropic example
if params.get("json_mode"):
    messages = self.format_history(history) + [
        {"role": "user", "content": [{"type": "text", "text": prompt}]}
    ]
    # Anthropic reads JSON from a tool call; simplified here

Parsing system prompts and conversation history

System prompts differ: OpenAI uses a `system` role, while Anthropic treats the first message as a system prompt if its role is `”assistant”` with a `”content”` field marked as `type: “system”`. The adapters implement `format_history` accordingly, preventing the **inconsistent conversation state** error when switching models.

Configuring per‑model parameters

The UI could expose sliders for `temperature` and `max_tokens`. The router forwards the exact dictionary; each adapter extracts the keys it supports, leaving unknown keys untouched.

Testing and debugging the live routing system

Simulating API errors and rate limiting

Wrap the generator in a try/except block inside `ModelRouter.send_prompt`. On `openai.RateLimitError` or `anthropic.RateLimitError`, put a descriptive token into the queue so the UI shows “Rate limit exceeded”.

try:
    for token in generator:
        self.response_queue.put(token)
except (openai.RateLimitError, anthropic.RateLimitError) as e:
    self.response_queue.put(f"\n[Error] {e}")

Logging prompts and responses

A lightweight logger writes each interaction to a rotating file. This creates an audit trail required for production use.

import logging
logger = logging.getLogger("router")
handler = logging.handlers.RotatingFileHandler(
    "router.log", maxBytes=1_048_576, backupCount=3
)
logger.setLevel(logging.INFO)
logger.addHandler(handler)

def log_interaction(provider, prompt, response):
    logger.info(
        f"{provider.upper()} | Prompt: {prompt!r} | Response: {response!r}"
    )

The `log_interaction` call can be placed after the response queue is emptied.

Validating model outputs after a switch

After a provider switch, the UI can display a small banner summarizing the current conversation length for that provider. This lets the developer verify that the history count matches expectations.

def show_history_len(self):
    hist = self.router.providers[self.router.active]["history"]
    self.output.configure(state="normal")
    self.output.insert("end", f"\n[Info] {self.router.active} history length: {len(hist)}\n")
    self.output.configure(state="disabled")

Trade‑offs & when not to use this

AspectTkinter desktop routerWeb‑based selector
LatencyMinimal, no network round‑trip for UI assetsDependent on HTTP server and browser
Multi‑user supportSingle user per processNaturally multi‑user
DistributionPackaged as an executable with PyInstallerRequires hosting infrastructure
UI richnessLimited widgets, less stylingFull HTML/CSS/JS capabilities
ComplexitySimple thread‑based designNeeds async server, auth, scaling

**When to choose the desktop router** – you need a quick, locally controlled tool for model comparison, you do not require simultaneous access by many users, and you prefer a single binary that can run on a developer workstation.

**When to avoid it** – you must share the interface across a team, need real‑time collaboration, or want to embed the router inside another web service.

Common errors and fixes

`TclError: main thread is not in main loop`

*Cause* – Attempting to modify a Tkinter widget from a background thread.

*Fix* – Always route UI updates through `after` or a thread‑safe queue. The `poll_queue` pattern above satisfies this rule.

Inconsistent conversation state after switching models

*Cause* – The new provider receives a history that contains roles it does not understand (e.g., OpenAI `system` messages sent to Claude).

*Fix* – Each adapter’s `format_history` must translate roles to the provider’s accepted schema. The code in the adapters already performs this conversion.

Rate‑limit exception crashes the GUI

*Cause* – Unhandled exception propagates out of the worker thread.

*Fix* – Catch provider‑specific rate‑limit errors inside `ModelRouter.send_prompt` and push an error token to the queue as shown earlier.

Frequently asked questions

How do I prevent the Tkinter GUI from freezing when calling an AI API?

Run the API call inside a `threading.Thread`. Return token fragments through a `queue.Queue` and let the main thread pull them with `after` for safe UI updates.

Can I compare GPT‑4 and Claude‑Opus outputs side‑by‑side?

Yes. Extend the UI with two read‑only text widgets and modify the router to dispatch the same prompt to both adapters simultaneously. Each worker thread writes to its own queue, which the corresponding output pane polls.

What is the best way to package this dashboard for non‑technical users?

Use `pyinstaller –onefile –windowed dashboard.py`. The generated executable bundles the Python interpreter, required wheels, and Tkinter runtime, allowing distribution without a separate Python installation.

How can I add a local LLM served by Ollama to the router?

Create a new adapter that posts to `http://localhost:11434/api/chat` using `requests`. Implement `generate` as a generator that yields `response[“message”][“content”]` chunks, then register the adapter in `ModelRouter.providers`.

Is it safe to store API keys in the code repository?<
Written by

’m Nilesh, a Software Development Engineer with 2+ years of experience, specializing in Go, JavaScript, Python, Docker, Kubernetes, Git, Jenkins, microservices, and system design (LLD/HLD), backed by a strong foundation in data structures and algorithms. Alongside my engineering journey, I bring 4+ years of hands-on experience in SEO, where I’ve worked extensively on content strategy, keyword research, technical SEO, and organic growth, helping products and businesses scale efficiently by aligning solid technology with search-driven performance.