This guide is for Python developers building automated image annotation pipelines with Anthropic Claude. You will construct a desktop agent that processes image folders, generates structured bounding box data, and exports to standard computer vision formats.

How automatic image annotation with Claude 3.5 works

Automatic image annotation uses multimodal large language models to detect objects and return coordinate data without manual labeling. You send images to the Anthropic API with structured prompts requesting specific output formats, parse the JSON responses, validate coordinates against image dimensions, and convert results to Pascal VOC XML or COCO JSON for training downstream models.

Claude 3.5 Sonnet accepts base64-encoded images and can follow explicit instructions to return normalized bounding box coordinates or pixel values. This enables rapid prototyping of annotation workflows without training custom detection models.

Prerequisites

  • Python 3.10 or higher
  • Anthropic API key with access to Claude 3.5 Sonnet
  • anthropic>=0.28.0 (Python SDK documentation)
  • Pillow>=10.0.0 for image handling
  • opencv-python>=4.8.0 for validation and preview
  • lxml>=4.9.0 for XML export

Install dependencies:

python -m pip install anthropic>=0.28.0 Pillow>=10.0.0 opencv-python>=4.8.0 lxml>=4.9.0

Configure your API key:

export ANTHROPIC_API_KEY='your-key-here'

Architecture of an AI driven annotation agent

An agentic annotation system requires three coordinated components: an event loop managing API interactions, a transformation layer parsing unstructured LLM outputs into strict formats, and a validation gate catching model hallucinations before they corrupt training data.

The event loop must handle asynchronous API calls without blocking the GUI thread. The transformation layer converts Claude’s text responses into structured annotation objects. The validation gate checks coordinate bounds against actual image dimensions.

flowchart LR
    A[Image Folder] --> B[Resize & Encode]
    B --> C[Anthropic API Call]
    C --> D[JSON Extraction]
    D --> E[Coordinate Validation]
    E --> F{Valid?}
    F -->|Yes| G[Pascal VOC / COCO Export]
    F -->|No| H[Retry with Hint]
    H --> C
    G --> I[GUI Preview]

Setting up the client and image encoding

High-resolution images exceed Claude’s input token limits and API latency thresholds. Resize images to a maximum dimension of 1568 pixels before encoding, maintaining aspect ratio to preserve detection accuracy.

Create encoder.py with image preprocessing:

import base64
import io
from pathlib import Path

from PIL import Image


MAX_DIMENSION = 1568


def encode_image_for_api(image_path: Path) -> tuple[str, tuple[int, int]]:
    """Return base64 string and final dimensions after resizing."""
    with Image.open(image_path) as img:
        width, height = img.size
        max_side = max(width, height)
        
        if max_side > MAX_DIMENSION:
            scale = MAX_DIMENSION / max_side
            new_width = int(width * scale)
            new_height = int(height * scale)
            img = img.resize((new_width, new_height), Image.Resampling.LANCZOS)
        else:
            new_width, new_height = width, height
        
        buffer = io.BytesIO()
        img_format = "PNG" if img.mode in ("RGBA", "P") else "JPEG"
        img = img.convert("RGB") if img.mode in ("RGBA", "P", "CMYK") else img
        img.save(buffer, format=img_format)
        
        encoded = base64.b64encode(buffer.getvalue()).decode("utf-8")
        return encoded, (new_width, new_height)

The function returns both the encoded string and final dimensions. Store the original dimensions separately if you need to scale coordinates back for full-resolution export.

Building the annotation logic with Claude 3.5

Claude does not have a native annotation output format. You must construct explicit system prompts that constrain the model to return parseable JSON with normalized coordinates (0.0 to 1.0).

Create annotator.py:

import json
import os
from pathlib import Path

from anthropic import Anthropic, APIError, APITimeoutError

from encoder import encode_image_for_api


SYSTEM_PROMPT = """You are a computer vision annotation assistant. Analyze the image and detect all requested object categories.

Respond ONLY with a JSON object in this exact format:
{
  "annotations": [
    {
      "label": "category_name",
      "bbox": [x_min, y_min, x_max, y_max]
    }
  ]
}

Bounding box coordinates must be normalized between 0.0 and 1.0 relative to image dimensions.
x_min and y_min are the top-left corner. x_max and y_max are the bottom-right corner.
Do not include any text outside the JSON."""


class AnnotationError(Exception):
    pass


class ClaudeAnnotator:
    def __init__(self, api_key: str | None = None):
        self.client = Anthropic(api_key=api_key or os.environ.get("ANTHROPIC_API_KEY"))
        self.model = "claude-3-5-sonnet-20241022"
    
    def annotate_image(
        self,
        image_path: Path,
        categories: list[str],
        max_retries: int = 2
    ) -> list[dict]:
        """Return validated annotations for a single image."""
        encoded_image, (width, height) = encode_image_for_api(image_path)
        
        user_content = f"""Detect these categories: {', '.join(categories)}.

Image (base64 JPEG):
{encoded_image[:100]}...[{len(encoded_image)} chars]"""
        
        for attempt in range(max_retries):
            try:
                message = self.client.messages.create(
                    model=self.model,
                    max_tokens=4096,
                    system=SYSTEM_PROMPT,
                    messages=[{"role": "user", "content": user_content}],
                    temperature=0.0,
                )
                
                raw_text = message.content[0].text
                annotations = self._parse_response(raw_text)
                validated = self._validate_coordinates(annotations, width, height)
                
                if not validated and attempt < max_retries - 1:
                    user_content += "\n\nPrevious response contained invalid coordinates. Ensure all values are between 0.0 and 1.0."
                    continue
                
                return validated
                
            except (APIError, APITimeoutError) as e:
                if attempt == max_retries - 1:
                    raise AnnotationError(f"API failed after {max_retries} attempts: {e}")
                continue
        
        return []
    
    def _parse_response(self, text: str) -> list[dict]:
        """Extract JSON from model response."""
        text = text.strip()
        if "```json" in text:
            text = text.split("```json")[1].split("```")[0]
        elif "```" in text:
            text = text.split("```")[1].split("```")[0]
        
        try:
            data = json.loads(text)
            return data.get("annotations", [])
        except json.JSONDecodeError as e:
            raise AnnotationError(f"Failed to parse JSON: {e}")
    
    def _validate_coordinates(
        self,
        annotations: list[dict],
        width: int,
        height: int
    ) -> list[dict]:
        """Filter out annotations with invalid coordinates."""
        valid = []
        
        for ann in annotations:
            bbox = ann.get("bbox")
            if not bbox or len(bbox) != 4:
                continue
            
            x_min, y_min, x_max, y_max = bbox
            
            # Check for negative or inverted boxes
            if not all(0.0 <= c <= 1.0 for c in bbox):
                continue
            if x_min >= x_max or y_min >= y_max:
                continue
            
            # Convert to pixel coordinates for export
            valid.append({
                "label": ann.get("label", "unknown"),
                "bbox_norm": [x_min, y_min, x_max, y_max],
                "bbox_pix": [
                    int(x_min * width),
                    int(y_min * height),
                    int(x_max * width),
                    int(y_max * height)
                ],
                "width": width,
                "height": height
            })
        
        return valid

The validation layer catches the most common failure mode: hallucinated coordinates outside image bounds. The retry mechanism appends explicit correction instructions when validation fails.

Creating a GUI wrapper for batch processing

Tkinter provides sufficient functionality for folder selection and progress display without adding heavy dependencies. The critical requirement is threading: API calls must run on a background thread to prevent the GUI from freezing.

Create gui.py:

import threading
import tkinter as tk
from pathlib import Path
from tkinter import filedialog, messagebox, ttk

from annotator import AnnotationError, ClaudeAnnotator
from exporter import export_pascal_voc, export_coco_json


class AnnotationGUI:
    def __init__(self):
        self.root = tk.Tk()
        self.root.title("Claude Image Annotator")
        self.root.geometry("600x400")
        
        self.annotator = ClaudeAnnotator()
        self.running = False
        
        self._build_ui()
    
    def _build_ui(self):
        # Input selection
        tk.Label(self.root, text="Image Folder:").pack(anchor="w", padx=10, pady=(10, 0))
        self.folder_var = tk.StringVar()
        tk.Entry(self.root, textvariable=self.folder_var, width=50).pack(padx=10)
        tk.Button(self.root, text="Browse...", command=self._browse_folder).pack(anchor="w", padx=10)
        
        # Categories
        tk.Label(self.root, text="Categories (comma-separated):").pack(anchor="w", padx=10, pady=(10, 0))
        self.cat_var = tk.StringVar(value="person, car, dog, cat")
        tk.Entry(self.root, textvariable=self.cat_var, width=50).pack(padx=10)
        
        # Output format
        tk.Label(self.root, text="Output Format:").pack(anchor="w", padx=10, pady=(10, 0))
        self.format_var = tk.StringVar(value="pascal_voc")
        tk.Radiobutton(self.root, text="Pascal VOC XML", variable=self.format_var, value="pascal_voc").pack(anchor="w", padx=20)
        tk.Radiobutton(self.root, text="COCO JSON", variable=self.format_var, value="coco").pack(anchor="w", padx=20)
        
        # Progress
        self.progress = ttk.Progressbar(self.root, mode="determinate")
        self.progress.pack(fill="x", padx=10, pady=20)
        
        self.status_var = tk.StringVar(value="Ready")
        tk.Label(self.root, textvariable=self.status_var).pack()
        
        # Buttons
        btn_frame = tk.Frame(self.root)
        btn_frame.pack(pady=10)
        self.start_btn = tk.Button(btn_frame, text="Start Annotation", command=self._start_annotation)
        self.start_btn.pack(side="left", padx=5)
        tk.Button(btn_frame, text="Quit", command=self.root.quit).pack(side="left", padx=5)
    
    def _browse_folder(self):
        folder = filedialog.askdirectory()
        if folder:
            self.folder_var.set(folder)
    
    def _start_annotation(self):
        if self.running:
            return
        
        folder = Path(self.folder_var.get())
        if not folder.exists():
            messagebox.showerror("Error", "Please select a valid folder")
            return
        
        categories = [c.strip() for c in self.cat_var.get().split(",") if c.strip()]
        if not categories:
            messagebox.showerror("Error", "Please specify at least one category")
            return
        
        self.running = True
        self.start_btn.config(state="disabled")
        self.status_var.set("Starting...")
        
        # Run annotation in background thread
        thread = threading.Thread(
            target=self._annotation_worker,
            args=(folder, categories, self.format_var.get())
        )
        thread.daemon = True
        thread.start()
    
    def _annotation_worker(self, folder: Path, categories: list[str], fmt: str):
        """Background thread: process images and update GUI via after()."""
        image_files = list(folder.glob("*.jpg")) + list(folder.glob("*.jpeg")) + list(folder.glob("*.png"))
        total = len(image_files)
        
        results = []
        
        for i, img_path in enumerate(image_files, 1):
            self.root.after(0, lambda i=i, t=total: self._update_progress(i, t, img_path.name))
            
            try:
                annotations = self.annotator.annotate_image(img_path, categories)
                results.append({
                    "file_name": img_path.name,
                    "path": img_path,
                    "annotations": annotations
                })
            except AnnotationError as e:
                self.root.after(0, lambda e=e: self._log_error(img_path.name, str(e)))
        
        # Export results
        output_dir = folder / "annotations"
        output_dir.mkdir(exist_ok=True)
        
        if fmt == "pascal_voc":
            for r in results:
                export_pascal_voc(r, output_dir)
        else:
            export_coco_json(results, output_dir / "annotations.json")
        
        self.root.after(0, self._annotation_complete, len(results), str(output_dir))
    
    def _update_progress(self, current: int, total: int, name: str):
        self.progress["maximum"] = total
        self.progress["value"] = current
        self.status_var.set(f"Processing {name} ({current}/{total})")
    
    def _log_error(self, filename: str, error: str):
        # Simple console logging; could extend to GUI log widget
        print(f"Error processing {filename}: {error}")
    
    def _annotation_complete(self, count: int, output_dir: str):
        self.running = False
        self.start_btn.config(state="normal")
        self.status_var.set(f"Complete: {count} images processed")
        messagebox.showinfo("Complete", f"Annotations saved to:\n{output_dir}")
    
    def run(self):
        self.root.mainloop()


if __name__ == "__main__":
    app = AnnotationGUI()
    app.run()

The threading.Thread wrapping prevents the GUI from freezing during 5-15 second API calls. The worker thread communicates with the main thread only through root.after() for thread-safe GUI updates.

Exporting to standard annotation formats

Create exporter.py for format conversion:

import json
import xml.etree.ElementTree as ET
from pathlib import Path

from lxml import etree


def export_pascal_voc(result: dict, output_dir: Path) -> None:
    """Write single image annotations to Pascal VOC XML."""
    root = ET.Element("annotation")
    
    ET.SubElement(root, "folder").text = "images"
    ET.SubElement(root, "filename").text = result["file_name"]
    ET.SubElement(root, "path").text = str(result["path"])
    
    size = ET.SubElement(root, "size")
    if result["annotations"]:
        h, w = result["annotations"][0]["height"], result["annotations"][0]["width"]
    else:
        h, w = 0, 0
    ET.SubElement(size, "width").text = str(w)
    ET.SubElement(size, "height").text = str(h)
    ET.SubElement(size, "depth").text = "3"
    
    ET.SubElement(root, "segmented").text = "0"
    
    for ann in result["annotations"]:
        obj = ET.SubElement(root, "object")
        ET.SubElement(obj, "name").text = ann["label"]
        ET.SubElement(obj, "pose").text = "Unspecified"
        ET.SubElement(obj, "truncated").text = "0"
        ET.SubElement(obj, "difficult").text = "0"
        
        bbox = ET.SubElement(obj, "bndbox")
        x_min, y_min, x_max, y_max = ann["bbox_pix"]
        ET.SubElement(bbox, "xmin").text = str(x_min)
        ET.SubElement(bbox, "ymin").text = str(y_min)
        ET.SubElement(bbox, "xmax").text = str(x_max)
        ET.SubElement(bbox, "ymax").text = str(y_max)
    
    # Pretty print with lxml
    tree = etree.ElementTree(root)
    output_path = output_dir / f"{result['file_name'].rsplit('.', 1)[0]}.xml"
    tree.write(output_path, pretty_print=True, xml_declaration=True, encoding="utf-8")


def export_coco_json(results: list[dict], output_path: Path) -> None:
    """Write batch results to COCO JSON format."""
    images = []
    annotations = []
    categories = {}
    ann_id = 1
    
    for img_id, result in enumerate(results, 1):
        if not result["annotations"]:
            continue
        
        dims = result["annotations"][0]
        images.append({
            "id": img_id,
            "file_name": result["file_name"],
            "width": dims["width"],
            "height": dims["height"]
        })
        
        for ann in result["annotations"]:
            label = ann["label"]
            if label not in categories:
                categories[label] = len(categories) + 1
            
            x_min, y_min, x_max, y_max = ann["bbox_pix"]
            width = x_max - x_min
            height = y_max - y_min
            
            annotations.append({
                "id": ann_id,
                "image_id": img_id,
                "category_id": categories[label],
                "bbox": [x_min, y_min, width, height],
                "area": width * height,
                "iscrowd": 0
            })
            ann_id += 1
    
    coco = {
        "images": images,
        "annotations": annotations,
        "categories": [{"id": v, "name": k} for k, v in categories.items()]
    }
    
    output_path.write_text(json.dumps(coco, indent=2), encoding="utf-8")

Accuracy vs Speed: when to use LLM annotation

ApproachSetup CostPer-Image CostLatencyAccuracyBest For
Claude 3.5 SonnetAPI key only$0.003-0.0153-15sModerate (zero-shot)Rapid prototyping, novel categories
Fine-tuned YOLOv8Hours of training + labeled dataCompute only10-50msHighProduction, fixed category sets
Manual LabelImgInstallationLabor hoursHuman speed
Written by

’m Nilesh, a Software Development Engineer with 2+ years of experience, specializing in Go, JavaScript, Python, Docker, Kubernetes, Git, Jenkins, microservices, and system design (LLD/HLD), backed by a strong foundation in data structures and algorithms. Alongside my engineering journey, I bring 4+ years of hands-on experience in SEO, where I’ve worked extensively on content strategy, keyword research, technical SEO, and organic growth, helping products and businesses scale efficiently by aligning solid technology with search-driven performance.