Building a pluggable CLI lets you invoke DALL‑E, Stable Diffusion, or Midjourney from a single command line without rewriting prompt handling each time. The guide below walks senior Python developers through a clean architecture, real‑world error handling, and packaging for distribution.
How to build a multi‑model AI image generator CLI
A pluggable CLI defines a **plugin interface** that each model adapter implements. The core program discovers available adapters at runtime, validates a unified set of arguments, and forwards the request to the selected adapter. Users switch generators with a `–model` flag while all other flags stay the same, achieving a single, consistent experience across disparate APIs.
Prerequisites
- Python 3.10 or newer
- `pip` or `poetry` for dependency management
- OpenAI account with an API key (`OPENAI_API_KEY`)
- Optional: a GPU with ≥8 GB VRAM for local Stable Diffusion inference
- Packages (install with `poetry add` or `pip install`): `typer`, `pydantic`, `openai`, `diffusers[torch]`, `replicate`, `midjourney-api`, `importlib-metadata`
Core architecture & concepts
The system consists of three layers:
- **CLI orchestrator** – parses arguments with Typer, discovers plugins, and forwards the request.
- **Plugin interface** – an abstract base class that enforces `generate` method signature and common configuration.
- **Model adapters** – concrete implementations for DALL‑E, Stable Diffusion, and Midjourney.
graph LR
A[CLI entry point] --> B[Typer parser]
B --> C[Plugin manager]
C --> D[Selected adapter]
D --> E[Model‑specific API call]
E --> F[Image bytes / URL]
F --> G[Save to disk]
The diagram shows a linear flow: the user runs the CLI, Typer extracts flags, the manager loads the appropriate adapter, the adapter talks to its backend, and the result is written locally.
Defining the plugin interface
Create a reusable contract in `ai_image/plugins/base.py`. The interface uses `abc.ABC` and Pydantic for typed settings.
# ai_image/plugins/base.py
from abc import ABC, abstractmethod
from pydantic import BaseModel, Field
from pathlib import Path
from typing import Tuple
class ModelConfig(BaseModel):
"""Common configuration fields exposed to the CLI."""
prompt: str = Field(..., description="Text prompt for image generation")
width: int = Field(512, ge=256, le=2048, description="Desired image width")
height: int = Field(512, ge=256, le=2048, description="Desired image height")
output_dir: Path = Field(Path.cwd() / "generated", description="Directory for saved images")
seed: int | None = Field(None, description="Random seed for reproducibility")
class ImageGenerator(ABC):
"""Abstract base that every model plugin must subclass."""
@abstractmethod
def generate(self, cfg: ModelConfig) -> Tuple[bytes, str]:
"""
Produce an image.
Returns:
tuple(image_bytes, file_extension) – e.g. (b'...', 'png')
"""
raise NotImplementedError
The `ModelConfig` object guarantees that every plugin receives the same fields, while each implementation can map them to provider‑specific parameters.
Implementing the DALL‑E plugin
Create `ai_image/plugins/dalle.py`. The OpenAI Python SDK provides `openai.Image.create`.
# ai_image/plugins/dalle.py
import os
import openai
from .base import ImageGenerator, ModelConfig
class DalleGenerator(ImageGenerator):
"""Adapter for OpenAI's DALL‑E image generation endpoint."""
def __init__(self) -> None:
api_key = os.getenv("OPENAI_API_KEY")
if not api_key:
raise EnvironmentError("OPENAI_API_KEY environment variable is missing")
openai.api_key = api_key
def generate(self, cfg: ModelConfig) -> tuple[bytes, str]:
try:
response = openai.Image.create(
prompt=cfg.prompt,
n=1,
size=f"{cfg.width}x{cfg.height}",
response_format="b64_json",
)
except openai.error.RateLimitError as exc:
raise RuntimeError("OpenAI rate limit exceeded") from exc
except openai.error.InvalidRequestError as exc:
raise RuntimeError(f"OpenAI rejected the request: {exc.user_message}") from exc
b64_image = response["data"][0]["b64_json"]
return (base64.b64decode(b64_image), "png")
The adapter validates the API key at construction time and translates width/height into the `size` string expected by DALL‑E. Errors are re‑raised with concise messages for the orchestrator to display.
Implementing the Stable Diffusion plugin
The local path uses `diffusers` and optionally a remote Replicate endpoint.
# ai_image/plugins/stable_diffusion.py
import torch
from diffusers import StableDiffusionPipeline
from .base import ImageGenerator, ModelConfig
class StableDiffusionLocal(ImageGenerator):
"""Runs Stable Diffusion locally via the diffusers pipeline."""
def __init__(self, model_id: str = "runwayml/stable-diffusion-v1-5"):
self.pipeline = StableDiffusionPipeline.from_pretrained(
model_id,
torch_dtype=torch.float16,
revision="fp16",
).to("cuda")
self.pipeline.safety_checker = None # disable for speed; handle policy elsewhere
def generate(self, cfg: ModelConfig) -> tuple[bytes, str]:
generator = torch.Generator(device="cuda").manual_seed(cfg.seed) if cfg.seed else None
image = self.pipeline(
prompt=cfg.prompt,
height=cfg.height,
width=cfg.width,
generator=generator,
).images[0]
buffer = BytesIO()
image.save(buffer, format="PNG")
return (buffer.getvalue(), "png")
If a user cannot run a GPU, they can swap the class with a Replicate wrapper that calls `replicate.run`. Both adapters respect the same `ModelConfig`.
Implementing the Midjourney pseudo‑plugin
Midjourney does not expose an official API. The unofficial `midjourney-api` library automates the Discord bot workflow.
# ai_image/plugins/midjourney.py
import os
from midjourney_api import MidjourneyClient
from .base import ImageGenerator, ModelConfig
class MidjourneyGenerator(ImageGenerator):
"""Wraps the unofficial Midjourney Discord bot via midjourney-api."""
def __init__(self):
token = os.getenv("MIDJOURNEY_TOKEN")
if not token:
raise EnvironmentError("MIDJOURNEY_TOKEN environment variable is missing")
self.client = MidjourneyClient(token)
def generate(self, cfg: ModelConfig) -> tuple[bytes, str]:
try:
result = self.client.imagine(prompt=cfg.prompt, width=cfg.width, height=cfg.height)
except Exception as exc:
raise RuntimeError("Midjourney request failed") from exc
# `result.image_bytes` holds raw PNG data
return (result.image_bytes, "png")
The wrapper handles Discord long‑polling internally; the CLI only needs to catch generic exceptions and present a clear message. Note the legal caveat in the FAQ.
Building the core CLI orchestrator
The orchestrator lives in `ai_image/main.py`. Typer provides a concise API, and `importlib.metadata.entry_points` enables dynamic plugin discovery.
# ai_image/main.py
import importlib
from pathlib import Path
from typing import Callable
import typer
from ai_image.plugins.base import ModelConfig, ImageGenerator
app = typer.Typer(help="Unified AI image generation CLI")
# Register plugins via entry points in pyproject.toml
def load_plugins() -> dict[str, Callable[[], ImageGenerator]]:
plugins: dict[str, Callable[[], ImageGenerator]] = {}
for ep in importlib.metadata.entry_points(group="ai_image.plugins"):
plugins[ep.name] = ep.load
return plugins
plugins = load_plugins()
available_models = ", ".join(plugins.keys())
@app.command()
def generate(
model: str = typer.Option(..., help=f"Select model ({available_models})"),
prompt: str = typer.Option(..., help="Text prompt for image generation"),
width: int = typer.Option(512, help="Image width in pixels"),
height: int = typer.Option(512, help="Image height in pixels"),
output_dir: Path = typer.Option(Path.cwd() / "generated", help="Where to save images"),
seed: int = typer.Option(None, help="Random seed for reproducibility"),
):
"""
Generate an image using the selected AI model.
"""
if model not in plugins:
typer.echo(f"Model '{model}' not registered. Available: {available_models}")
raise typer.Exit(code=1)
cfg = ModelConfig(
prompt=prompt,
width=width,
height=height,
output_dir=output_dir,
seed=seed,
)
generator = plugins[model]()
image_bytes, ext = generator.generate(cfg)
output_dir.mkdir(parents=True, exist_ok=True)
out_path = output_dir / f"{model}_{seed or 'rand'}.{ext}"
out_path.write_bytes(image_bytes)
typer.echo(f"Image saved to {out_path}")
if __name__ == "__main__":
app()
**Key points**
- **Dynamic discovery**: plugins register through the `ai_image.plugins` entry‑point group. Adding a new model only requires publishing a package with the entry point; no CLI code changes.
- **Shared arguments**: `ModelConfig` ensures every plugin sees the same flag set.
- **Error propagation**: plugins raise `RuntimeError` or `EnvironmentError`; the orchestrator prints the message and exits with a non‑zero status.
Registering plugins (pyproject.toml excerpt)
[project.entry-points."ai_image.plugins"]
dalle = "ai_image.plugins.dalle:DalleGenerator"
stable_diffusion = "ai_image.plugins.stable_diffusion:StableDiffusionLocal"
midjourney = "ai_image.plugins.midjourney:MidjourneyGenerator"
Advanced features and configuration
Caching API responses
Repeated prompts can consume credits. A lightweight file‑based cache reduces cost.
# ai_image/cache.py
import hashlib
import json
from pathlib import Path
CACHE_ROOT = Path.home() / ".ai_image_cache"
CACHE_ROOT.mkdir(exist_ok=True)
def cache_key(cfg: ModelConfig) -> str:
payload = json.dumps(cfg.dict(), sort_keys=True).encode()
return hashlib.sha256(payload).hexdigest()
def get_cached(cfg: ModelConfig) -> bytes | None:
path = CACHE_ROOT / f"{cache_key(cfg)}.png"
return path.read_bytes() if path.is_file() else None
def store_cached(cfg: ModelConfig, data: bytes) -> None:
(CACHE_ROOT / f"{cache_key(cfg)}.png").write_bytes(data)
Integrate the cache in `generate` before invoking the plugin; if a hit occurs, skip the remote call.
Custom output naming scheme
Allow users to control naming via a `–name-template` flag that supports `{model}`, `{prompt_hash}`, and `{timestamp}` placeholders. Implementation is straightforward string formatting inside the orchestrator.
Trade‑offs & when not to use this
| Aspect | Unified CLI (plugin layer) | Direct model‑specific script |
|---|---|---|
| User experience | Single command set, easy model switching | Separate scripts, each with its own flags |
| Access to model‑specific features | Limited to what the abstract interface exposes | Full access to every provider‑specific parameter |
| Development overhead | Requires plugin scaffolding and entry‑point management | Minimal initial code |
| Maintenance | One place to update error handling, logging, caching | Duplicate logic across scripts |
**When to avoid the unified approach**
- Your project depends on advanced model‑specific controls (e.g., DALL‑E’s `style` or Midjourney’s `–v 5` flags) that cannot be expressed in the common `ModelConfig`.
- Deployment targets have strict memory limits and cannot host the Stable Diffusion plugin locally; a lightweight single‑model script may be preferable.
Common errors and fixes
`EnvironmentError: OPENAI_API_KEY environment variable is missing`
**Cause**: The key is not exported in the shell. **Fix**: Add the variable to your session or a `.env` file and load it with `python-dotenv`.
export OPENAI_API_KEY=sk-...
# or
echo "OPENAI_API_KEY=sk-..." >> .env
`torch.cuda.OutOfMemoryError: CUDA out of memory`
**Cause**: The GPU cannot accommodate the requested image size or batch. **Fix**: Reduce `–width`/`–height` or switch to the `fp16` checkpoint (already used). For CPUs, instantiate the pipeline with `torch_dtype=torch.float32` and `device_map=”auto”`.
self.pipeline = StableDiffusionPipeline.from_pretrained(
model_id,
torch_dtype=torch.float32,
device_map="auto",
)
`Midjourney request failed` (generic exception)
**Cause**: Discord token expired or the wrapper library is out of sync with Discord changes. **Fix**: Regenerate the Discord user token, update `midjourney-api` to the latest version, and verify the bot is still invited to the server.
`OpenAI rate limit exceeded`
**Cause**: Too many requests in a short interval. **Fix**: Implement exponential backoff. See the article on [Retry and Backoff Strategy for AI APIs: 5 Tips (2026)](https://nileshblog.tech/?p=6770) for a reusable decorator.
Frequently asked questions
Can I use this CLI to run Stable Diffusion locally without an internet connection?
Yes. After the model checkpoint is downloaded once via the `diffusers` library, all subsequent generations run entirely offline.
How do I handle different image aspect ratios between DALL‑E and Stable Diffusion?
The unified `ModelConfig` includes `width` and `height`. Each plugin maps those fields to its native API – DALL‑E receives a `size` string, while Stable Diffusion passes the explicit dimensions to the pipeline. Post‑generation resizing can be added if needed.
Is automating Midjourney via a CLI against their terms of service?
Using unofficial wrappers to control the Discord bot likely violates Midjourney’s Terms of Service. The approach is suitable for personal experimentation but should not be used in production or commercial settings.
Wrap‑up
A pluggable CLI abstracts away the quirks of DALL‑E, Stable Diffusion, and Midjourney while keeping each adapter isolated. By defining a strict `ModelConfig` and leveraging Python entry points, you gain a single command line that can evolve as new generators appear. For projects that need full access to every provider‑specific flag, a dedicated script per model remains the simpler choice.