This guide is for Python developers building automated image annotation pipelines with Anthropic Claude. You will construct a desktop agent that processes image folders, generates structured bounding box data, and exports to standard computer vision formats.
How automatic image annotation with Claude 3.5 works
Automatic image annotation uses multimodal large language models to detect objects and return coordinate data without manual labeling. You send images to the Anthropic API with structured prompts requesting specific output formats, parse the JSON responses, validate coordinates against image dimensions, and convert results to Pascal VOC XML or COCO JSON for training downstream models.
Claude 3.5 Sonnet accepts base64-encoded images and can follow explicit instructions to return normalized bounding box coordinates or pixel values. This enables rapid prototyping of annotation workflows without training custom detection models.
Prerequisites
- Python 3.10 or higher
- Anthropic API key with access to Claude 3.5 Sonnet
anthropic>=0.28.0(Python SDK documentation)Pillow>=10.0.0for image handlingopencv-python>=4.8.0for validation and previewlxml>=4.9.0for XML export
Install dependencies:
python -m pip install anthropic>=0.28.0 Pillow>=10.0.0 opencv-python>=4.8.0 lxml>=4.9.0
Configure your API key:
export ANTHROPIC_API_KEY='your-key-here'
Architecture of an AI driven annotation agent
An agentic annotation system requires three coordinated components: an event loop managing API interactions, a transformation layer parsing unstructured LLM outputs into strict formats, and a validation gate catching model hallucinations before they corrupt training data.
The event loop must handle asynchronous API calls without blocking the GUI thread. The transformation layer converts Claude’s text responses into structured annotation objects. The validation gate checks coordinate bounds against actual image dimensions.
flowchart LR
A[Image Folder] --> B[Resize & Encode]
B --> C[Anthropic API Call]
C --> D[JSON Extraction]
D --> E[Coordinate Validation]
E --> F{Valid?}
F -->|Yes| G[Pascal VOC / COCO Export]
F -->|No| H[Retry with Hint]
H --> C
G --> I[GUI Preview]
Setting up the client and image encoding
High-resolution images exceed Claude’s input token limits and API latency thresholds. Resize images to a maximum dimension of 1568 pixels before encoding, maintaining aspect ratio to preserve detection accuracy.
Create encoder.py with image preprocessing:
import base64
import io
from pathlib import Path
from PIL import Image
MAX_DIMENSION = 1568
def encode_image_for_api(image_path: Path) -> tuple[str, tuple[int, int]]:
"""Return base64 string and final dimensions after resizing."""
with Image.open(image_path) as img:
width, height = img.size
max_side = max(width, height)
if max_side > MAX_DIMENSION:
scale = MAX_DIMENSION / max_side
new_width = int(width * scale)
new_height = int(height * scale)
img = img.resize((new_width, new_height), Image.Resampling.LANCZOS)
else:
new_width, new_height = width, height
buffer = io.BytesIO()
img_format = "PNG" if img.mode in ("RGBA", "P") else "JPEG"
img = img.convert("RGB") if img.mode in ("RGBA", "P", "CMYK") else img
img.save(buffer, format=img_format)
encoded = base64.b64encode(buffer.getvalue()).decode("utf-8")
return encoded, (new_width, new_height)
The function returns both the encoded string and final dimensions. Store the original dimensions separately if you need to scale coordinates back for full-resolution export.
Building the annotation logic with Claude 3.5
Claude does not have a native annotation output format. You must construct explicit system prompts that constrain the model to return parseable JSON with normalized coordinates (0.0 to 1.0).
Create annotator.py:
import json
import os
from pathlib import Path
from anthropic import Anthropic, APIError, APITimeoutError
from encoder import encode_image_for_api
SYSTEM_PROMPT = """You are a computer vision annotation assistant. Analyze the image and detect all requested object categories.
Respond ONLY with a JSON object in this exact format:
{
"annotations": [
{
"label": "category_name",
"bbox": [x_min, y_min, x_max, y_max]
}
]
}
Bounding box coordinates must be normalized between 0.0 and 1.0 relative to image dimensions.
x_min and y_min are the top-left corner. x_max and y_max are the bottom-right corner.
Do not include any text outside the JSON."""
class AnnotationError(Exception):
pass
class ClaudeAnnotator:
def __init__(self, api_key: str | None = None):
self.client = Anthropic(api_key=api_key or os.environ.get("ANTHROPIC_API_KEY"))
self.model = "claude-3-5-sonnet-20241022"
def annotate_image(
self,
image_path: Path,
categories: list[str],
max_retries: int = 2
) -> list[dict]:
"""Return validated annotations for a single image."""
encoded_image, (width, height) = encode_image_for_api(image_path)
user_content = f"""Detect these categories: {', '.join(categories)}.
Image (base64 JPEG):
{encoded_image[:100]}...[{len(encoded_image)} chars]"""
for attempt in range(max_retries):
try:
message = self.client.messages.create(
model=self.model,
max_tokens=4096,
system=SYSTEM_PROMPT,
messages=[{"role": "user", "content": user_content}],
temperature=0.0,
)
raw_text = message.content[0].text
annotations = self._parse_response(raw_text)
validated = self._validate_coordinates(annotations, width, height)
if not validated and attempt < max_retries - 1:
user_content += "\n\nPrevious response contained invalid coordinates. Ensure all values are between 0.0 and 1.0."
continue
return validated
except (APIError, APITimeoutError) as e:
if attempt == max_retries - 1:
raise AnnotationError(f"API failed after {max_retries} attempts: {e}")
continue
return []
def _parse_response(self, text: str) -> list[dict]:
"""Extract JSON from model response."""
text = text.strip()
if "```json" in text:
text = text.split("```json")[1].split("```")[0]
elif "```" in text:
text = text.split("```")[1].split("```")[0]
try:
data = json.loads(text)
return data.get("annotations", [])
except json.JSONDecodeError as e:
raise AnnotationError(f"Failed to parse JSON: {e}")
def _validate_coordinates(
self,
annotations: list[dict],
width: int,
height: int
) -> list[dict]:
"""Filter out annotations with invalid coordinates."""
valid = []
for ann in annotations:
bbox = ann.get("bbox")
if not bbox or len(bbox) != 4:
continue
x_min, y_min, x_max, y_max = bbox
# Check for negative or inverted boxes
if not all(0.0 <= c <= 1.0 for c in bbox):
continue
if x_min >= x_max or y_min >= y_max:
continue
# Convert to pixel coordinates for export
valid.append({
"label": ann.get("label", "unknown"),
"bbox_norm": [x_min, y_min, x_max, y_max],
"bbox_pix": [
int(x_min * width),
int(y_min * height),
int(x_max * width),
int(y_max * height)
],
"width": width,
"height": height
})
return valid
The validation layer catches the most common failure mode: hallucinated coordinates outside image bounds. The retry mechanism appends explicit correction instructions when validation fails.
Creating a GUI wrapper for batch processing
Tkinter provides sufficient functionality for folder selection and progress display without adding heavy dependencies. The critical requirement is threading: API calls must run on a background thread to prevent the GUI from freezing.
Create gui.py:
import threading
import tkinter as tk
from pathlib import Path
from tkinter import filedialog, messagebox, ttk
from annotator import AnnotationError, ClaudeAnnotator
from exporter import export_pascal_voc, export_coco_json
class AnnotationGUI:
def __init__(self):
self.root = tk.Tk()
self.root.title("Claude Image Annotator")
self.root.geometry("600x400")
self.annotator = ClaudeAnnotator()
self.running = False
self._build_ui()
def _build_ui(self):
# Input selection
tk.Label(self.root, text="Image Folder:").pack(anchor="w", padx=10, pady=(10, 0))
self.folder_var = tk.StringVar()
tk.Entry(self.root, textvariable=self.folder_var, width=50).pack(padx=10)
tk.Button(self.root, text="Browse...", command=self._browse_folder).pack(anchor="w", padx=10)
# Categories
tk.Label(self.root, text="Categories (comma-separated):").pack(anchor="w", padx=10, pady=(10, 0))
self.cat_var = tk.StringVar(value="person, car, dog, cat")
tk.Entry(self.root, textvariable=self.cat_var, width=50).pack(padx=10)
# Output format
tk.Label(self.root, text="Output Format:").pack(anchor="w", padx=10, pady=(10, 0))
self.format_var = tk.StringVar(value="pascal_voc")
tk.Radiobutton(self.root, text="Pascal VOC XML", variable=self.format_var, value="pascal_voc").pack(anchor="w", padx=20)
tk.Radiobutton(self.root, text="COCO JSON", variable=self.format_var, value="coco").pack(anchor="w", padx=20)
# Progress
self.progress = ttk.Progressbar(self.root, mode="determinate")
self.progress.pack(fill="x", padx=10, pady=20)
self.status_var = tk.StringVar(value="Ready")
tk.Label(self.root, textvariable=self.status_var).pack()
# Buttons
btn_frame = tk.Frame(self.root)
btn_frame.pack(pady=10)
self.start_btn = tk.Button(btn_frame, text="Start Annotation", command=self._start_annotation)
self.start_btn.pack(side="left", padx=5)
tk.Button(btn_frame, text="Quit", command=self.root.quit).pack(side="left", padx=5)
def _browse_folder(self):
folder = filedialog.askdirectory()
if folder:
self.folder_var.set(folder)
def _start_annotation(self):
if self.running:
return
folder = Path(self.folder_var.get())
if not folder.exists():
messagebox.showerror("Error", "Please select a valid folder")
return
categories = [c.strip() for c in self.cat_var.get().split(",") if c.strip()]
if not categories:
messagebox.showerror("Error", "Please specify at least one category")
return
self.running = True
self.start_btn.config(state="disabled")
self.status_var.set("Starting...")
# Run annotation in background thread
thread = threading.Thread(
target=self._annotation_worker,
args=(folder, categories, self.format_var.get())
)
thread.daemon = True
thread.start()
def _annotation_worker(self, folder: Path, categories: list[str], fmt: str):
"""Background thread: process images and update GUI via after()."""
image_files = list(folder.glob("*.jpg")) + list(folder.glob("*.jpeg")) + list(folder.glob("*.png"))
total = len(image_files)
results = []
for i, img_path in enumerate(image_files, 1):
self.root.after(0, lambda i=i, t=total: self._update_progress(i, t, img_path.name))
try:
annotations = self.annotator.annotate_image(img_path, categories)
results.append({
"file_name": img_path.name,
"path": img_path,
"annotations": annotations
})
except AnnotationError as e:
self.root.after(0, lambda e=e: self._log_error(img_path.name, str(e)))
# Export results
output_dir = folder / "annotations"
output_dir.mkdir(exist_ok=True)
if fmt == "pascal_voc":
for r in results:
export_pascal_voc(r, output_dir)
else:
export_coco_json(results, output_dir / "annotations.json")
self.root.after(0, self._annotation_complete, len(results), str(output_dir))
def _update_progress(self, current: int, total: int, name: str):
self.progress["maximum"] = total
self.progress["value"] = current
self.status_var.set(f"Processing {name} ({current}/{total})")
def _log_error(self, filename: str, error: str):
# Simple console logging; could extend to GUI log widget
print(f"Error processing {filename}: {error}")
def _annotation_complete(self, count: int, output_dir: str):
self.running = False
self.start_btn.config(state="normal")
self.status_var.set(f"Complete: {count} images processed")
messagebox.showinfo("Complete", f"Annotations saved to:\n{output_dir}")
def run(self):
self.root.mainloop()
if __name__ == "__main__":
app = AnnotationGUI()
app.run()
The threading.Thread wrapping prevents the GUI from freezing during 5-15 second API calls. The worker thread communicates with the main thread only through root.after() for thread-safe GUI updates.
Exporting to standard annotation formats
Create exporter.py for format conversion:
import json
import xml.etree.ElementTree as ET
from pathlib import Path
from lxml import etree
def export_pascal_voc(result: dict, output_dir: Path) -> None:
"""Write single image annotations to Pascal VOC XML."""
root = ET.Element("annotation")
ET.SubElement(root, "folder").text = "images"
ET.SubElement(root, "filename").text = result["file_name"]
ET.SubElement(root, "path").text = str(result["path"])
size = ET.SubElement(root, "size")
if result["annotations"]:
h, w = result["annotations"][0]["height"], result["annotations"][0]["width"]
else:
h, w = 0, 0
ET.SubElement(size, "width").text = str(w)
ET.SubElement(size, "height").text = str(h)
ET.SubElement(size, "depth").text = "3"
ET.SubElement(root, "segmented").text = "0"
for ann in result["annotations"]:
obj = ET.SubElement(root, "object")
ET.SubElement(obj, "name").text = ann["label"]
ET.SubElement(obj, "pose").text = "Unspecified"
ET.SubElement(obj, "truncated").text = "0"
ET.SubElement(obj, "difficult").text = "0"
bbox = ET.SubElement(obj, "bndbox")
x_min, y_min, x_max, y_max = ann["bbox_pix"]
ET.SubElement(bbox, "xmin").text = str(x_min)
ET.SubElement(bbox, "ymin").text = str(y_min)
ET.SubElement(bbox, "xmax").text = str(x_max)
ET.SubElement(bbox, "ymax").text = str(y_max)
# Pretty print with lxml
tree = etree.ElementTree(root)
output_path = output_dir / f"{result['file_name'].rsplit('.', 1)[0]}.xml"
tree.write(output_path, pretty_print=True, xml_declaration=True, encoding="utf-8")
def export_coco_json(results: list[dict], output_path: Path) -> None:
"""Write batch results to COCO JSON format."""
images = []
annotations = []
categories = {}
ann_id = 1
for img_id, result in enumerate(results, 1):
if not result["annotations"]:
continue
dims = result["annotations"][0]
images.append({
"id": img_id,
"file_name": result["file_name"],
"width": dims["width"],
"height": dims["height"]
})
for ann in result["annotations"]:
label = ann["label"]
if label not in categories:
categories[label] = len(categories) + 1
x_min, y_min, x_max, y_max = ann["bbox_pix"]
width = x_max - x_min
height = y_max - y_min
annotations.append({
"id": ann_id,
"image_id": img_id,
"category_id": categories[label],
"bbox": [x_min, y_min, width, height],
"area": width * height,
"iscrowd": 0
})
ann_id += 1
coco = {
"images": images,
"annotations": annotations,
"categories": [{"id": v, "name": k} for k, v in categories.items()]
}
output_path.write_text(json.dumps(coco, indent=2), encoding="utf-8")
Accuracy vs Speed: when to use LLM annotation
| Approach | Setup Cost | Per-Image Cost | Latency | Accuracy | Best For |
|---|---|---|---|---|---|
| Claude 3.5 Sonnet | API key only | $0.003-0.015 | 3-15s | Moderate (zero-shot) | Rapid prototyping, novel categories |
| Fine-tuned YOLOv8 | Hours of training + labeled data | Compute only | 10-50ms | High | Production, fixed category sets |
| Manual LabelImg | Installation | Labor hours | Human speed |