When a floor associate asks, “Do we have size 10 black sneakers in aisle 5?” the response must be instantaneous, accurate, and never expose the shop’s internal APIs to the public internet. I hit this wall while prototyping a voice‑driven inventory checker: my Streamlit demo was sending the Claude API key and the ChromaDB token straight from the browser. The moment a curious shopper opened the dev tools, the secrets were visible. The fix? Split the UI from the credential‑bearing agent and run everything that talks to Claude or the vector store behind a hardened FastAPI service.
- Separate the GUI (Streamlit/PyQt) from a secure FastAPI backend that holds all secrets.
- Use Anthropic’s tool‑calling to let Claude fetch inventory vectors on demand.
- Stream partial responses back to the UI for a natural “thinking…” feel.
- Guard the backend with rate‑limit, retry, and circuit‑breaker logic.
- Docker‑compose the two services and use env‑file secrets for production.
Before you start: Python 3.11+, `anthropic>=0.30.0`, `fastapi`, `uvicorn`, `streamlit` **or** `PyQt6`, `chromadb` (or Pinecone), OpenAI Whisper API key (or Vosk for offline STT), and a Claude 3.5 Sonnet API key. Use a virtual environment and keep all secrets in a `.env` file.
Secure Voice AI for Retail: Build a RAG‑Enabled Inventory Lookup (2026 Guide)
A secure voice AI for retail uses a RAG‑enabled Claude AI agent. You build a GUI with Streamlit/PyQt for voice input. The key is separating the interface from a secure backend service. This backend, holding your API/DB keys, performs all inventory vector searches and LLM calls, preventing credential exposure while providing real‑time stock answers.
System architecture & workflow
The diagram below shows the split‑brain design. The front‑end never sees a secret; it only posts audio blobs (or transcribed text) to `/query`. The FastAPI backend owns the Claude client, the vector DB, and the tool‑calling loop. Responses stream back as Server‑Sent Events (SSE) so the UI can render a “typing…” animation.
flowchart LR
UI[UI (Streamlit / PyQt)] -->|POST /query| API[FastAPI Backend]
API -->|tool call| Claude[Claude SDK]
API -->|vector search| VDB[Vector DB (Chroma/Pinecone)]
Claude -->|tool result| API
API -->|SSE streaming| UI
1. Building the GUI frontend
You can pick either Streamlit for quick web‑style prototyping or PyQt6 for a native desktop experience. The code below works for both; the only difference is the event loop integration.
1.1 Install the UI stack
# Streamlit
pip install streamlit==1.38.0
# PyQt6 (optional)
pip install PyQt6==6.7.0
1.2 Voice input component
For a web UI we can leverage the browser’s Web Speech API via Streamlit’s `st.audio` and JavaScript shim. For a desktop UI we’ll call Whisper’s REST endpoint from the backend (kept secret) and display the transcription.
# ui.py – shared logic
import os
import httpx
import streamlit as st # replace with PyQt6 widgets if you go native
API_URL = os.getenv("BACKEND_URL", "http://localhost:8000")
def transcribe_and_send(audio_bytes: bytes):
"""Send raw audio to the backend; returns a request ID."""
resp = httpx.post(f"{API_URL}/query", files={"audio": ("voice.wav", audio_bytes)})
resp.raise_for_status()
return resp.json()["request_id"]
def stream_response(request_id: str):
"""Yield server‑sent events from the backend."""
with httpx.stream("GET", f"{API_URL}/stream/{request_id}") as s:
for line in s.iter_lines():
if line:
yield line.decode()
# Streamlit UI
st.title("Retail Voice Assistant")
audio = st.audio_input("Speak your request", format="wav")
if audio:
request_id = transcribe_and_send(audio.read())
placeholder = st.empty()
for chunk in stream_response(request_id):
placeholder.markdown(chunk, unsafe_allow_html=True)
**My take:** Streamlit’s `st.audio_input` is still experimental (v1.38). If you need rock‑solid support across browsers, drop a tiny React front‑end that forwards the Blob to the FastAPI endpoint and keep Streamlit for the chat view only.
1.3 Session state handling
# keep a list of exchanged messages per user session
if "history" not in st.session_state:
st.session_state.history = []
def append_message(role: str, content: str):
st.session_state.history.append({"role": role, "content": content})
st.write(f"**{role.title()}:** {content}")
2. Connecting to the Claude AI Agent backend
The backend owns the Claude SDK, the vector DB client, and the tool‑calling loop. We’ll expose two endpoints:
- `POST /query` – receives audio, runs Whisper (or Vosk), then hands the transcript to Claude.
- `GET /stream/{request_id}` – streams partial Claude responses as SSE.
2.1 FastAPI skeleton
# backend.py
import os
import uuid
import asyncio
from fastapi import FastAPI, File, UploadFile, HTTPException
from fastapi.responses import StreamingResponse
from anthropic import Anthropic, AsyncClient
import chromadb
from chromadb.config import Settings
app = FastAPI()
anthropic = AsyncClient(api_key=os.getenv("ANTHROPIC_API_KEY"))
vector_db = chromadb.Client(Settings(chroma_db_impl="duckdb+parquet", persist_directory="./chroma"))
collection = vector_db.get_or_create_collection(name="inventory")
# Simple in‑memory store for pending streams
STREAMS = {}
@app.post("/query")
async def query(audio: UploadFile = File(...)):
# 1️⃣ Transcribe – keep the Whisper call inside the backend
transcript = await transcribe(audio)
request_id = str(uuid.uuid4())
# 2️⃣ Kick off the agent loop in background
asyncio.create_task(agent_loop(request_id, transcript))
return {"request_id": request_id}
@app.get("/stream/{request_id}")
def stream(request_id: str):
if request_id not in STREAMS:
raise HTTPException(404, "No such request")
async def event_generator():
async for chunk in STREAMS[request_id]:
yield f"data: {chunk}\n\n"
# clean up
del STREAMS[request_id]
return StreamingResponse(event_generator(), media_type="text/event-stream")
2.2 Whisper transcription helper
# helper.py
import httpx
import os
WHISPER_ENDPOINT = "https://api.openai.com/v1/audio/transcriptions"
async def transcribe(upload: UploadFile) -> str:
headers = {"Authorization": f"Bearer {os.getenv('OPENAI_API_KEY')}"}
files = {"file": (upload.filename, await upload.read(), upload.content_type)}
async with httpx.AsyncClient() as client:
resp = await client.post(WHISPER_ENDPOINT, headers=headers, files=files, timeout=30)
resp.raise_for_status()
return resp.json()["text"]
**Tip:** If you prefer an offline model, swap the `transcribe` function for a Vosk inference pipeline. The only change is the import and model loading; the rest of the backend stays identical.
2.3 RAG tool definition
Claude’s tool‑calling lets the LLM ask the backend to run a vector search. We expose a single tool: `search_inventory(query: str) -> str`.
# rag_tool.py
from typing import List, Dict
def embed_documents(docs: List[Dict]) -> None:
# called once during bootstrap to load product catalog
vectors = [{"id": d["sku"], "embedding": d["embedding"], "metadata": d} for d in docs]
collection.add(
documents=[d["description"] for d in docs],
ids=[d["sku"] for d in docs],
embeddings=[d["embedding"] for d in docs],
metadatas=[d["metadata"] for d in docs],
)
def search_inventory(query: str, top_k: int = 5) -> str:
results = collection.query(query_texts=[query], n_results=top_k)
# Format a concise answer for Claude
msgs = []
for i, doc in enumerate(results["documents"][0]):
sku = results["ids"][0][i]
msgs.append(f"- SKU {sku}: {doc}")
return "\n".join(msgs) if msgs else "No matching items."
2.4 The secure agent loop
# agent_loop.py
import json
import asyncio
from anthropic import AsyncClient
from rag_tool import search_inventory
MAX_RETRIES = 3
RETRY_BACKOFF = 1.5 # seconds
async def agent_loop(request_id: str, user_query: str):
# streaming channel for this request
queue = asyncio.Queue()
STREAMS[request_id] = _queue_to_async_iter(queue)
# Build the initial message payload
messages = [{"role": "user", "content": user_query}]
client = AsyncClient(api_key=os.getenv("ANTHROPIC_API_KEY"))
for attempt in range(1, MAX_RETRIES + 1):
try:
async with client.messages.stream(
model="claude-3-5-sonnet-20240620",
max_tokens=1024,
temperature=0.0,
messages=messages,
tools=[{
"name": "search_inventory",
"description": "Search product catalog",
"input_schema": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"]
}
}]
) as stream:
async for event in stream:
if event.type == "content_block_start":
continue
if event.type == "content_block_delta":
await queue.put(event.delta)
if event.type == "tool_use":
# Claude asks for a tool call
query = event.input["query"]
result = search_inventory(query)
# Append tool result and let Claude continue
messages.append({"role": "assistant", "content": None, "tool_use_id": event.id})
messages.append({"role": "tool", "content": result, "tool_use_id": event.id})
if event.type == "message_stop":
await queue.put("[DONE]")
return
except Exception as exc:
if attempt == MAX_RETRIES:
await queue.put(f"[ERROR] {type(exc).__name__}: {exc}")
return
await asyncio.sleep(RETRY_BACKOFF ** attempt)
# utility to convert Queue to async iterator
async def _queue_to_async_iter(queue: asyncio.Queue):
while True:
item = await queue.get()
if item == "[DONE]":
break
yield item
3. Implementing RAG without exposing keys
3.1 Private vector DB setup
# bootstrap.py – run once to ingest catalog
import json, os
from rag_tool import embed_documents
with open("catalog.json") as f:
products = json.load(f) # each entry: {sku, description, embedding, metadata...}
embed_documents(products)
print("✅ Catalog indexed")
The embeddings are produced ahead of time with `sentence‑transformers==3.0.1` or a vendor‑specific service. Store them in a local DuckDB file (`chroma` default) so the backend never reaches out to a remote vector API at query time.
3.2 Why keys never leave the server
All calls that need the Claude API key, the OpenAI Whisper key, or the vector‑DB credentials happen **inside** the FastAPI process. The front‑end only sees an opaque `request_id`. Even if an attacker inspects network traffic, the only thing they can replay is the audio payload, not the secrets.
4. Handling edge cases & production gotchas
4.1 Network timeouts & retry logic
We already wrapped the Claude streaming call in an exponential‑backoff loop. For Whisper and vector DB calls, apply the same pattern:
# generic async retry decorator
def retry(times: int = 3, backoff: float = 1.2):
def decorator(fn):
async def wrapper(*args, **kwargs):
delay = backoff
for attempt in range(times):
try:
return await fn(*args, **kwargs)
except Exception as e:
if attempt == times - 1:
raise
await asyncio.sleep(delay)
delay *= backoff
return wrapper
return decorator
4.2 Rate limiting tool use
Claude’s pricing model charges per token, but you also risk 429 errors if the UI spams requests. Use an in‑memory token bucket per client IP.
from collections import defaultdict
import time
RATE_LIMIT = 5 # requests per minute
buckets = defaultdict(lambda: [0, time.time()]) # [count, reset_ts]
def allow_request(ip: str) -> bool:
count, reset = buckets[ip]
now = time.time()
if now > reset + 60:
buckets[ip] = [1, now]
return True
if count < RATE_LIMIT:
buckets[ip][0] += 1
return True
return False
Add this guard at the start of `/query`. If you need a more sophisticated solution, read my guide on **[Secure Multi‑Tenant API Keys & Prompts for Google ADK (2026)](https://nileshblog.tech/google-adk-secret-management/)**.
4.3 Async UI without freezing
Both Streamlit and PyQt support asyncio natively now. In Streamlit, use `st.experimental_async` (v1.38) to launch the background task without blocking the UI thread. In PyQt, spin up a `QThreadPool` and post results via signals.
# PyQt example snippet
from PyQt6.QtCore import QRunnable, QThreadPool, pyqtSignal, QObject
class WorkerSignals(QObject):
result = pyqtSignal(str)
class AgentWorker(QRunnable):
def __init__(self, audio_bytes):
super().__init__()
self.audio = audio_bytes
self.signals = WorkerSignals()
def run(self):
# same logic as FastAPI client, but local call
loop = asyncio.new_event_loop()
result = loop.run_until_complete(transcribe_and_send(self.audio))
self.signals.result.emit(result)
# usage
pool = QThreadPool.globalInstance()
worker = AgentWorker(audio_blob)
worker.signals.result.connect(handle_result)
pool.start(worker)
5. Deployment: secure hosting
5.1 Docker‑compose the two services
# docker-compose.yml
version: "3.9"
services:
backend:
build: ./backend
env_file: .env
ports: ["8000: