Building a pluggable CLI lets you invoke DALL‑E, Stable Diffusion, or Midjourney from a single command line without rewriting prompt handling each time. The guide below walks senior Python developers through a clean architecture, real‑world error handling, and packaging for distribution.

How to build a multi‑model AI image generator CLI

A pluggable CLI defines a **plugin interface** that each model adapter implements. The core program discovers available adapters at runtime, validates a unified set of arguments, and forwards the request to the selected adapter. Users switch generators with a `–model` flag while all other flags stay the same, achieving a single, consistent experience across disparate APIs.

Prerequisites

  • Python 3.10 or newer
  • `pip` or `poetry` for dependency management
  • OpenAI account with an API key (`OPENAI_API_KEY`)
  • Optional: a GPU with ≥8 GB VRAM for local Stable Diffusion inference
  • Packages (install with `poetry add` or `pip install`): `typer`, `pydantic`, `openai`, `diffusers[torch]`, `replicate`, `midjourney-api`, `importlib-metadata`

Core architecture & concepts

The system consists of three layers:

  1. **CLI orchestrator** – parses arguments with Typer, discovers plugins, and forwards the request.
  2. **Plugin interface** – an abstract base class that enforces `generate` method signature and common configuration.
  3. **Model adapters** – concrete implementations for DALL‑E, Stable Diffusion, and Midjourney.
graph LR
    A[CLI entry point] --> B[Typer parser]
    B --> C[Plugin manager]
    C --> D[Selected adapter]
    D --> E[Model‑specific API call]
    E --> F[Image bytes / URL]
    F --> G[Save to disk]

The diagram shows a linear flow: the user runs the CLI, Typer extracts flags, the manager loads the appropriate adapter, the adapter talks to its backend, and the result is written locally.

Defining the plugin interface

Create a reusable contract in `ai_image/plugins/base.py`. The interface uses `abc.ABC` and Pydantic for typed settings.

# ai_image/plugins/base.py
from abc import ABC, abstractmethod
from pydantic import BaseModel, Field
from pathlib import Path
from typing import Tuple

class ModelConfig(BaseModel):
    """Common configuration fields exposed to the CLI."""
    prompt: str = Field(..., description="Text prompt for image generation")
    width: int = Field(512, ge=256, le=2048, description="Desired image width")
    height: int = Field(512, ge=256, le=2048, description="Desired image height")
    output_dir: Path = Field(Path.cwd() / "generated", description="Directory for saved images")
    seed: int | None = Field(None, description="Random seed for reproducibility")

class ImageGenerator(ABC):
    """Abstract base that every model plugin must subclass."""
    
    @abstractmethod
    def generate(self, cfg: ModelConfig) -> Tuple[bytes, str]:
        """
        Produce an image.

        Returns:
            tuple(image_bytes, file_extension) – e.g. (b'...', 'png')
        """
        raise NotImplementedError

The `ModelConfig` object guarantees that every plugin receives the same fields, while each implementation can map them to provider‑specific parameters.

Implementing the DALL‑E plugin

Create `ai_image/plugins/dalle.py`. The OpenAI Python SDK provides `openai.Image.create`.

# ai_image/plugins/dalle.py
import os
import openai
from .base import ImageGenerator, ModelConfig

class DalleGenerator(ImageGenerator):
    """Adapter for OpenAI's DALL‑E image generation endpoint."""

    def __init__(self) -> None:
        api_key = os.getenv("OPENAI_API_KEY")
        if not api_key:
            raise EnvironmentError("OPENAI_API_KEY environment variable is missing")
        openai.api_key = api_key

    def generate(self, cfg: ModelConfig) -> tuple[bytes, str]:
        try:
            response = openai.Image.create(
                prompt=cfg.prompt,
                n=1,
                size=f"{cfg.width}x{cfg.height}",
                response_format="b64_json",
            )
        except openai.error.RateLimitError as exc:
            raise RuntimeError("OpenAI rate limit exceeded") from exc
        except openai.error.InvalidRequestError as exc:
            raise RuntimeError(f"OpenAI rejected the request: {exc.user_message}") from exc

        b64_image = response["data"][0]["b64_json"]
        return (base64.b64decode(b64_image), "png")

The adapter validates the API key at construction time and translates width/height into the `size` string expected by DALL‑E. Errors are re‑raised with concise messages for the orchestrator to display.

Implementing the Stable Diffusion plugin

The local path uses `diffusers` and optionally a remote Replicate endpoint.

# ai_image/plugins/stable_diffusion.py
import torch
from diffusers import StableDiffusionPipeline
from .base import ImageGenerator, ModelConfig

class StableDiffusionLocal(ImageGenerator):
    """Runs Stable Diffusion locally via the diffusers pipeline."""

    def __init__(self, model_id: str = "runwayml/stable-diffusion-v1-5"):
        self.pipeline = StableDiffusionPipeline.from_pretrained(
            model_id,
            torch_dtype=torch.float16,
            revision="fp16",
        ).to("cuda")
        self.pipeline.safety_checker = None  # disable for speed; handle policy elsewhere

    def generate(self, cfg: ModelConfig) -> tuple[bytes, str]:
        generator = torch.Generator(device="cuda").manual_seed(cfg.seed) if cfg.seed else None
        image = self.pipeline(
            prompt=cfg.prompt,
            height=cfg.height,
            width=cfg.width,
            generator=generator,
        ).images[0]
        buffer = BytesIO()
        image.save(buffer, format="PNG")
        return (buffer.getvalue(), "png")

If a user cannot run a GPU, they can swap the class with a Replicate wrapper that calls `replicate.run`. Both adapters respect the same `ModelConfig`.

Implementing the Midjourney pseudo‑plugin

Midjourney does not expose an official API. The unofficial `midjourney-api` library automates the Discord bot workflow.

# ai_image/plugins/midjourney.py
import os
from midjourney_api import MidjourneyClient
from .base import ImageGenerator, ModelConfig

class MidjourneyGenerator(ImageGenerator):
    """Wraps the unofficial Midjourney Discord bot via midjourney-api."""

    def __init__(self):
        token = os.getenv("MIDJOURNEY_TOKEN")
        if not token:
            raise EnvironmentError("MIDJOURNEY_TOKEN environment variable is missing")
        self.client = MidjourneyClient(token)

    def generate(self, cfg: ModelConfig) -> tuple[bytes, str]:
        try:
            result = self.client.imagine(prompt=cfg.prompt, width=cfg.width, height=cfg.height)
        except Exception as exc:
            raise RuntimeError("Midjourney request failed") from exc
        # `result.image_bytes` holds raw PNG data
        return (result.image_bytes, "png")

The wrapper handles Discord long‑polling internally; the CLI only needs to catch generic exceptions and present a clear message. Note the legal caveat in the FAQ.

Building the core CLI orchestrator

The orchestrator lives in `ai_image/main.py`. Typer provides a concise API, and `importlib.metadata.entry_points` enables dynamic plugin discovery.

# ai_image/main.py
import importlib
from pathlib import Path
from typing import Callable
import typer
from ai_image.plugins.base import ModelConfig, ImageGenerator

app = typer.Typer(help="Unified AI image generation CLI")

# Register plugins via entry points in pyproject.toml
def load_plugins() -> dict[str, Callable[[], ImageGenerator]]:
    plugins: dict[str, Callable[[], ImageGenerator]] = {}
    for ep in importlib.metadata.entry_points(group="ai_image.plugins"):
        plugins[ep.name] = ep.load
    return plugins

plugins = load_plugins()
available_models = ", ".join(plugins.keys())

@app.command()
def generate(
    model: str = typer.Option(..., help=f"Select model ({available_models})"),
    prompt: str = typer.Option(..., help="Text prompt for image generation"),
    width: int = typer.Option(512, help="Image width in pixels"),
    height: int = typer.Option(512, help="Image height in pixels"),
    output_dir: Path = typer.Option(Path.cwd() / "generated", help="Where to save images"),
    seed: int = typer.Option(None, help="Random seed for reproducibility"),
):
    """
    Generate an image using the selected AI model.
    """
    if model not in plugins:
        typer.echo(f"Model '{model}' not registered. Available: {available_models}")
        raise typer.Exit(code=1)

    cfg = ModelConfig(
        prompt=prompt,
        width=width,
        height=height,
        output_dir=output_dir,
        seed=seed,
    )
    generator = plugins[model]()
    image_bytes, ext = generator.generate(cfg)

    output_dir.mkdir(parents=True, exist_ok=True)
    out_path = output_dir / f"{model}_{seed or 'rand'}.{ext}"
    out_path.write_bytes(image_bytes)
    typer.echo(f"Image saved to {out_path}")

if __name__ == "__main__":
    app()

**Key points**

  • **Dynamic discovery**: plugins register through the `ai_image.plugins` entry‑point group. Adding a new model only requires publishing a package with the entry point; no CLI code changes.
  • **Shared arguments**: `ModelConfig` ensures every plugin sees the same flag set.
  • **Error propagation**: plugins raise `RuntimeError` or `EnvironmentError`; the orchestrator prints the message and exits with a non‑zero status.

Registering plugins (pyproject.toml excerpt)

[project.entry-points."ai_image.plugins"]
dalle = "ai_image.plugins.dalle:DalleGenerator"
stable_diffusion = "ai_image.plugins.stable_diffusion:StableDiffusionLocal"
midjourney = "ai_image.plugins.midjourney:MidjourneyGenerator"

Advanced features and configuration

Caching API responses

Repeated prompts can consume credits. A lightweight file‑based cache reduces cost.

# ai_image/cache.py
import hashlib
import json
from pathlib import Path

CACHE_ROOT = Path.home() / ".ai_image_cache"
CACHE_ROOT.mkdir(exist_ok=True)

def cache_key(cfg: ModelConfig) -> str:
    payload = json.dumps(cfg.dict(), sort_keys=True).encode()
    return hashlib.sha256(payload).hexdigest()

def get_cached(cfg: ModelConfig) -> bytes | None:
    path = CACHE_ROOT / f"{cache_key(cfg)}.png"
    return path.read_bytes() if path.is_file() else None

def store_cached(cfg: ModelConfig, data: bytes) -> None:
    (CACHE_ROOT / f"{cache_key(cfg)}.png").write_bytes(data)

Integrate the cache in `generate` before invoking the plugin; if a hit occurs, skip the remote call.

Custom output naming scheme

Allow users to control naming via a `–name-template` flag that supports `{model}`, `{prompt_hash}`, and `{timestamp}` placeholders. Implementation is straightforward string formatting inside the orchestrator.

Trade‑offs & when not to use this

AspectUnified CLI (plugin layer)Direct model‑specific script
User experienceSingle command set, easy model switchingSeparate scripts, each with its own flags
Access to model‑specific featuresLimited to what the abstract interface exposesFull access to every provider‑specific parameter
Development overheadRequires plugin scaffolding and entry‑point managementMinimal initial code
MaintenanceOne place to update error handling, logging, cachingDuplicate logic across scripts

**When to avoid the unified approach**

  • Your project depends on advanced model‑specific controls (e.g., DALL‑E’s `style` or Midjourney’s `–v 5` flags) that cannot be expressed in the common `ModelConfig`.
  • Deployment targets have strict memory limits and cannot host the Stable Diffusion plugin locally; a lightweight single‑model script may be preferable.

Common errors and fixes

`EnvironmentError: OPENAI_API_KEY environment variable is missing`

**Cause**: The key is not exported in the shell. **Fix**: Add the variable to your session or a `.env` file and load it with `python-dotenv`.

export OPENAI_API_KEY=sk-...
# or
echo "OPENAI_API_KEY=sk-..." >> .env

`torch.cuda.OutOfMemoryError: CUDA out of memory`

**Cause**: The GPU cannot accommodate the requested image size or batch. **Fix**: Reduce `–width`/`–height` or switch to the `fp16` checkpoint (already used). For CPUs, instantiate the pipeline with `torch_dtype=torch.float32` and `device_map=”auto”`.

self.pipeline = StableDiffusionPipeline.from_pretrained(
    model_id,
    torch_dtype=torch.float32,
    device_map="auto",
)

`Midjourney request failed` (generic exception)

**Cause**: Discord token expired or the wrapper library is out of sync with Discord changes. **Fix**: Regenerate the Discord user token, update `midjourney-api` to the latest version, and verify the bot is still invited to the server.

`OpenAI rate limit exceeded`

**Cause**: Too many requests in a short interval. **Fix**: Implement exponential backoff. See the article on [Retry and Backoff Strategy for AI APIs: 5 Tips (2026)](https://nileshblog.tech/?p=6770) for a reusable decorator.

Frequently asked questions

Can I use this CLI to run Stable Diffusion locally without an internet connection?

Yes. After the model checkpoint is downloaded once via the `diffusers` library, all subsequent generations run entirely offline.

How do I handle different image aspect ratios between DALL‑E and Stable Diffusion?

The unified `ModelConfig` includes `width` and `height`. Each plugin maps those fields to its native API – DALL‑E receives a `size` string, while Stable Diffusion passes the explicit dimensions to the pipeline. Post‑generation resizing can be added if needed.

Is automating Midjourney via a CLI against their terms of service?

Using unofficial wrappers to control the Discord bot likely violates Midjourney’s Terms of Service. The approach is suitable for personal experimentation but should not be used in production or commercial settings.

Wrap‑up

A pluggable CLI abstracts away the quirks of DALL‑E, Stable Diffusion, and Midjourney while keeping each adapter isolated. By defining a strict `ModelConfig` and leveraging Python entry points, you gain a single command line that can evolve as new generators appear. For projects that need full access to every provider‑specific flag, a dedicated script per model remains the simpler choice.

Written by

’m Nilesh, a Software Development Engineer with 2+ years of experience, specializing in Go, JavaScript, Python, Docker, Kubernetes, Git, Jenkins, microservices, and system design (LLD/HLD), backed by a strong foundation in data structures and algorithms. Alongside my engineering journey, I bring 4+ years of hands-on experience in SEO, where I’ve worked extensively on content strategy, keyword research, technical SEO, and organic growth, helping products and businesses scale efficiently by aligning solid technology with search-driven performance.