When a floor associate asks, “Do we have size 10 black sneakers in aisle 5?” the response must be instantaneous, accurate, and never expose the shop’s internal APIs to the public internet. I hit this wall while prototyping a voice‑driven inventory checker: my Streamlit demo was sending the Claude API key and the ChromaDB token straight from the browser. The moment a curious shopper opened the dev tools, the secrets were visible. The fix? Split the UI from the credential‑bearing agent and run everything that talks to Claude or the vector store behind a hardened FastAPI service.

⚡ TL;DR — Key takeaways
  • Separate the GUI (Streamlit/PyQt) from a secure FastAPI backend that holds all secrets.
  • Use Anthropic’s tool‑calling to let Claude fetch inventory vectors on demand.
  • Stream partial responses back to the UI for a natural “thinking…” feel.
  • Guard the backend with rate‑limit, retry, and circuit‑breaker logic.
  • Docker‑compose the two services and use env‑file secrets for production.

Before you start: Python 3.11+, `anthropic>=0.30.0`, `fastapi`, `uvicorn`, `streamlit` **or** `PyQt6`, `chromadb` (or Pinecone), OpenAI Whisper API key (or Vosk for offline STT), and a Claude 3.5 Sonnet API key. Use a virtual environment and keep all secrets in a `.env` file.

Secure Voice AI for Retail: Build a RAG‑Enabled Inventory Lookup (2026 Guide)

A secure voice AI for retail uses a RAG‑enabled Claude AI agent. You build a GUI with Streamlit/PyQt for voice input. The key is separating the interface from a secure backend service. This backend, holding your API/DB keys, performs all inventory vector searches and LLM calls, preventing credential exposure while providing real‑time stock answers.

System architecture & workflow

The diagram below shows the split‑brain design. The front‑end never sees a secret; it only posts audio blobs (or transcribed text) to `/query`. The FastAPI backend owns the Claude client, the vector DB, and the tool‑calling loop. Responses stream back as Server‑Sent Events (SSE) so the UI can render a “typing…” animation.

flowchart LR
    UI[UI (Streamlit / PyQt)] -->|POST /query| API[FastAPI Backend]
    API -->|tool call| Claude[Claude SDK]
    API -->|vector search| VDB[Vector DB (Chroma/Pinecone)]
    Claude -->|tool result| API
    API -->|SSE streaming| UI

1. Building the GUI frontend

You can pick either Streamlit for quick web‑style prototyping or PyQt6 for a native desktop experience. The code below works for both; the only difference is the event loop integration.

1.1 Install the UI stack

# Streamlit
pip install streamlit==1.38.0

# PyQt6 (optional)
pip install PyQt6==6.7.0

1.2 Voice input component

For a web UI we can leverage the browser’s Web Speech API via Streamlit’s `st.audio` and JavaScript shim. For a desktop UI we’ll call Whisper’s REST endpoint from the backend (kept secret) and display the transcription.

# ui.py – shared logic
import os
import httpx
import streamlit as st  # replace with PyQt6 widgets if you go native

API_URL = os.getenv("BACKEND_URL", "http://localhost:8000")

def transcribe_and_send(audio_bytes: bytes):
    """Send raw audio to the backend; returns a request ID."""
    resp = httpx.post(f"{API_URL}/query", files={"audio": ("voice.wav", audio_bytes)})
    resp.raise_for_status()
    return resp.json()["request_id"]

def stream_response(request_id: str):
    """Yield server‑sent events from the backend."""
    with httpx.stream("GET", f"{API_URL}/stream/{request_id}") as s:
        for line in s.iter_lines():
            if line:
                yield line.decode()

# Streamlit UI
st.title("Retail Voice Assistant")
audio = st.audio_input("Speak your request", format="wav")
if audio:
    request_id = transcribe_and_send(audio.read())
    placeholder = st.empty()
    for chunk in stream_response(request_id):
        placeholder.markdown(chunk, unsafe_allow_html=True)

**My take:** Streamlit’s `st.audio_input` is still experimental (v1.38). If you need rock‑solid support across browsers, drop a tiny React front‑end that forwards the Blob to the FastAPI endpoint and keep Streamlit for the chat view only.

1.3 Session state handling

# keep a list of exchanged messages per user session
if "history" not in st.session_state:
    st.session_state.history = []

def append_message(role: str, content: str):
    st.session_state.history.append({"role": role, "content": content})
    st.write(f"**{role.title()}:** {content}")

2. Connecting to the Claude AI Agent backend

The backend owns the Claude SDK, the vector DB client, and the tool‑calling loop. We’ll expose two endpoints:

  • `POST /query` – receives audio, runs Whisper (or Vosk), then hands the transcript to Claude.
  • `GET /stream/{request_id}` – streams partial Claude responses as SSE.

2.1 FastAPI skeleton

# backend.py
import os
import uuid
import asyncio
from fastapi import FastAPI, File, UploadFile, HTTPException
from fastapi.responses import StreamingResponse
from anthropic import Anthropic, AsyncClient
import chromadb
from chromadb.config import Settings

app = FastAPI()
anthropic = AsyncClient(api_key=os.getenv("ANTHROPIC_API_KEY"))
vector_db = chromadb.Client(Settings(chroma_db_impl="duckdb+parquet", persist_directory="./chroma"))
collection = vector_db.get_or_create_collection(name="inventory")

# Simple in‑memory store for pending streams
STREAMS = {}

@app.post("/query")
async def query(audio: UploadFile = File(...)):
    # 1️⃣ Transcribe – keep the Whisper call inside the backend
    transcript = await transcribe(audio)
    request_id = str(uuid.uuid4())
    # 2️⃣ Kick off the agent loop in background
    asyncio.create_task(agent_loop(request_id, transcript))
    return {"request_id": request_id}

@app.get("/stream/{request_id}")
def stream(request_id: str):
    if request_id not in STREAMS:
        raise HTTPException(404, "No such request")
    async def event_generator():
        async for chunk in STREAMS[request_id]:
            yield f"data: {chunk}\n\n"
        # clean up
        del STREAMS[request_id]
    return StreamingResponse(event_generator(), media_type="text/event-stream")

2.2 Whisper transcription helper

# helper.py
import httpx
import os

WHISPER_ENDPOINT = "https://api.openai.com/v1/audio/transcriptions"

async def transcribe(upload: UploadFile) -> str:
    headers = {"Authorization": f"Bearer {os.getenv('OPENAI_API_KEY')}"}
    files = {"file": (upload.filename, await upload.read(), upload.content_type)}
    async with httpx.AsyncClient() as client:
        resp = await client.post(WHISPER_ENDPOINT, headers=headers, files=files, timeout=30)
    resp.raise_for_status()
    return resp.json()["text"]

**Tip:** If you prefer an offline model, swap the `transcribe` function for a Vosk inference pipeline. The only change is the import and model loading; the rest of the backend stays identical.

2.3 RAG tool definition

Claude’s tool‑calling lets the LLM ask the backend to run a vector search. We expose a single tool: `search_inventory(query: str) -> str`.

# rag_tool.py
from typing import List, Dict

def embed_documents(docs: List[Dict]) -> None:
    # called once during bootstrap to load product catalog
    vectors = [{"id": d["sku"], "embedding": d["embedding"], "metadata": d} for d in docs]
    collection.add(
        documents=[d["description"] for d in docs],
        ids=[d["sku"] for d in docs],
        embeddings=[d["embedding"] for d in docs],
        metadatas=[d["metadata"] for d in docs],
    )

def search_inventory(query: str, top_k: int = 5) -> str:
    results = collection.query(query_texts=[query], n_results=top_k)
    # Format a concise answer for Claude
    msgs = []
    for i, doc in enumerate(results["documents"][0]):
        sku = results["ids"][0][i]
        msgs.append(f"- SKU {sku}: {doc}")
    return "\n".join(msgs) if msgs else "No matching items."

2.4 The secure agent loop

# agent_loop.py
import json
import asyncio
from anthropic import AsyncClient
from rag_tool import search_inventory

MAX_RETRIES = 3
RETRY_BACKOFF = 1.5  # seconds

async def agent_loop(request_id: str, user_query: str):
    # streaming channel for this request
    queue = asyncio.Queue()
    STREAMS[request_id] = _queue_to_async_iter(queue)

    # Build the initial message payload
    messages = [{"role": "user", "content": user_query}]
    client = AsyncClient(api_key=os.getenv("ANTHROPIC_API_KEY"))

    for attempt in range(1, MAX_RETRIES + 1):
        try:
            async with client.messages.stream(
                model="claude-3-5-sonnet-20240620",
                max_tokens=1024,
                temperature=0.0,
                messages=messages,
                tools=[{
                    "name": "search_inventory",
                    "description": "Search product catalog",
                    "input_schema": {
                        "type": "object",
                        "properties": {"query": {"type": "string"}},
                        "required": ["query"]
                    }
                }]
            ) as stream:
                async for event in stream:
                    if event.type == "content_block_start":
                        continue
                    if event.type == "content_block_delta":
                        await queue.put(event.delta)
                    if event.type == "tool_use":
                        # Claude asks for a tool call
                        query = event.input["query"]
                        result = search_inventory(query)
                        # Append tool result and let Claude continue
                        messages.append({"role": "assistant", "content": None, "tool_use_id": event.id})
                        messages.append({"role": "tool", "content": result, "tool_use_id": event.id})
                    if event.type == "message_stop":
                        await queue.put("[DONE]")
                        return
        except Exception as exc:
            if attempt == MAX_RETRIES:
                await queue.put(f"[ERROR] {type(exc).__name__}: {exc}")
                return
            await asyncio.sleep(RETRY_BACKOFF ** attempt)
# utility to convert Queue to async iterator
async def _queue_to_async_iter(queue: asyncio.Queue):
    while True:
        item = await queue.get()
        if item == "[DONE]":
            break
        yield item

3. Implementing RAG without exposing keys

3.1 Private vector DB setup

# bootstrap.py – run once to ingest catalog
import json, os
from rag_tool import embed_documents

with open("catalog.json") as f:
    products = json.load(f)   # each entry: {sku, description, embedding, metadata...}
embed_documents(products)
print("✅ Catalog indexed")

The embeddings are produced ahead of time with `sentence‑transformers==3.0.1` or a vendor‑specific service. Store them in a local DuckDB file (`chroma` default) so the backend never reaches out to a remote vector API at query time.

3.2 Why keys never leave the server

All calls that need the Claude API key, the OpenAI Whisper key, or the vector‑DB credentials happen **inside** the FastAPI process. The front‑end only sees an opaque `request_id`. Even if an attacker inspects network traffic, the only thing they can replay is the audio payload, not the secrets.

4. Handling edge cases & production gotchas

4.1 Network timeouts & retry logic

We already wrapped the Claude streaming call in an exponential‑backoff loop. For Whisper and vector DB calls, apply the same pattern:

# generic async retry decorator
def retry(times: int = 3, backoff: float = 1.2):
    def decorator(fn):
        async def wrapper(*args, **kwargs):
            delay = backoff
            for attempt in range(times):
                try:
                    return await fn(*args, **kwargs)
                except Exception as e:
                    if attempt == times - 1:
                        raise
                    await asyncio.sleep(delay)
                    delay *= backoff
        return wrapper
    return decorator

4.2 Rate limiting tool use

Claude’s pricing model charges per token, but you also risk 429 errors if the UI spams requests. Use an in‑memory token bucket per client IP.

from collections import defaultdict
import time

RATE_LIMIT = 5  # requests per minute
buckets = defaultdict(lambda: [0, time.time()])  # [count, reset_ts]

def allow_request(ip: str) -> bool:
    count, reset = buckets[ip]
    now = time.time()
    if now > reset + 60:
        buckets[ip] = [1, now]
        return True
    if count < RATE_LIMIT:
        buckets[ip][0] += 1
        return True
    return False

Add this guard at the start of `/query`. If you need a more sophisticated solution, read my guide on **[Secure Multi‑Tenant API Keys & Prompts for Google ADK (2026)](https://nileshblog.tech/google-adk-secret-management/)**.

4.3 Async UI without freezing

Both Streamlit and PyQt support asyncio natively now. In Streamlit, use `st.experimental_async` (v1.38) to launch the background task without blocking the UI thread. In PyQt, spin up a `QThreadPool` and post results via signals.

# PyQt example snippet
from PyQt6.QtCore import QRunnable, QThreadPool, pyqtSignal, QObject

class WorkerSignals(QObject):
    result = pyqtSignal(str)

class AgentWorker(QRunnable):
    def __init__(self, audio_bytes):
        super().__init__()
        self.audio = audio_bytes
        self.signals = WorkerSignals()

    def run(self):
        # same logic as FastAPI client, but local call
        loop = asyncio.new_event_loop()
        result = loop.run_until_complete(transcribe_and_send(self.audio))
        self.signals.result.emit(result)

# usage
pool = QThreadPool.globalInstance()
worker = AgentWorker(audio_blob)
worker.signals.result.connect(handle_result)
pool.start(worker)

5. Deployment: secure hosting

5.1 Docker‑compose the two services

# docker-compose.yml
version: "3.9"
services:
  backend:
    build: ./backend
    env_file: .env
    ports: ["8000:
Written by

’m Nilesh, a Software Development Engineer with 2+ years of experience, specializing in Go, JavaScript, Python, Docker, Kubernetes, Git, Jenkins, microservices, and system design (LLD/HLD), backed by a strong foundation in data structures and algorithms. Alongside my engineering journey, I bring 4+ years of hands-on experience in SEO, where I’ve worked extensively on content strategy, keyword research, technical SEO, and organic growth, helping products and businesses scale efficiently by aligning solid technology with search-driven performance.