Skip to content

Repository files navigation

ArtSmoker

Smoke-testing your artwork!

Python FastAPI Amazon Bedrock License

ArtSmoker walkthrough — from a text prompt to production-ready 2D art to a fully-textured, game-engine-tailored 3D model

📌 0. Overview

ArtSmoker turns an idea into game-engine-tailored art — in minutes, with no pipeline for you to manage. Describe a character, prop, environment, or piece of key art in plain language and get production-ready 2D art, fully-textured 3D models, and video — all matched to your project's visual identity, all kept within your own environment. The newest AI image, editing, 3D, and video models sit behind one clean, artist-first interface with real creative controls: ArtSmoker runs the entire production pipeline for you, so your team directs the look instead of wrangling the machinery.

📝 The Problem

Creative teams in game and media studios want the leverage of generative AI — but today that power is locked behind developer tooling they were never meant to manage:

  • It's built for engineers, not artists — the best models live behind cloud consoles, command lines, SDKs, and REST APIs. No director or concept artist should need a terminal to make a piece of art.
  • Clear ideas, cryptic prompts — artists know exactly what they want, but the models don't take direction in plain creative language; consistent, on-brief results still hinge on prompt structure, negative prompts, and model-specific phrasing that sit between the brief and the output.
  • The best AI models are scattered and hard to run — powerful AI models for image, editing, 3D, and video ship constantly across different providers and formats; standing each one up (packaging, GPUs, quantization, scaling) is a full engineering project on its own.
  • Editing and 3D are separate worlds — inpainting, outpainting, recoloring, reference-guided edits, and turning a 2D concept into a textured 3D model each normally need their own tools, APIs, and specialists.
  • Staying on-brand is manual — keeping every asset true to your established look usually means babysitting each generation by hand.

📝 The Solution

ArtSmoker is a self-hosted creative studio that puts today's best generative models behind one artist-first interface — purpose-built for game asset production, and equally at home in film, advertising, e-commerce, publishing, and any team that lives on original visual content.

  • Describe it in plain language — ArtSmoker handles prompt decomposition, enhancement, and model-specific optimization behind the scenes. A guided Prompt Designer lets you shape each visual element — subject, scene, lighting, colour — with lock/vary controls to explore genuinely different directions without losing what's already working.
  • On-brand by default — feed ArtSmoker your existing art and its vision models learn your visual identity, so every asset comes out matching your project's look and feel.
  • 2D, edited, and in 3D — end to end — generate, then refine in place with inpainting, outpainting, recoloring, search-and-replace, and reference-guided edits; turn any 2D asset into a fully-textured, game-engine-tailored 3D model that drops straight into Unity, Unreal, or Blender — no manual modeling, UV unwrapping, or texture painting. Plus cinematic video and a multi-model chat studio for ideation.
  • Every model, one click — use the latest hosted models across regions, or deploy curated open-source models (Qwen-Image, FLUX.2, HunyuanImage, TripoSG, TRELLIS.2, and more) to your own GPUs with a single click — packaging, quantization, auto-scaling, and job tracking all handled, every model validated end-to-end before it ships.
  • Runs where you want — and your IP stays yours — install it on a single artist's desktop or a shared instance for the whole team; no GPU of your own required (the heavy compute runs on managed AWS services, or on auto-scaling endpoints ArtSmoker spins up and scales back to zero for you). It connects only to your own AWS account — artwork, prompts, styles, and generated assets stay in your environment, nothing goes to third-party services, and you keep full ownership of your creative IP.

Amazon Bedrock models: Claude Sonnet/Opus (prompt engineering & chat), Stable Diffusion 3.5 Large, Stable Image Ultra, Stable Image Core, Stability AI services (image editing), Nova Reel, Luma AI Ray (video generation), and 80+ LLMs from 16 providers for Chat Studio. Self-hosted models: Qwen-Image (text-to-image) & Qwen-Image-Edit (reference-guided + instruction editing, Apache-2.0), HunyuanImage 3.0 (BF16/NF4), FLUX.2, FLUX.1, TripoSG & TRELLIS.2 (image-to-3D), and more via Amazon SageMaker — with an extensible catalog for adding new models.

Get started now — jump to Prerequisites & Installation ▸

Language / 言語 / 语言 / 언어 / हिन्दी / Язык / Langue / Idioma

ArtSmoker supports 9 languages. Switch the UI language using the language buttons in the top navigation bar (EN | 日 | 中 | 한 | हिं | РУ | FR | ES | DE). Your selection is saved automatically.

Language README
English This document
日本語 (Japanese) README.ja.md
中文 (Chinese) README.zh.md
한국어 (Korean) README.ko.md
हिन्दी (Hindi) README.hi.md
Русский (Russian) README.ru.md
Français (French) README.fr.md
Español (Spanish) README.es.md
Deutsch (German) README.de.md

Multi-lingual prompt support:

  • Non-English prompts are automatically detected (Japanese, Chinese, Korean, Hindi, Russian, French, Spanish, and more) and translated to English before generation
  • A bilingual preview appears in the prompt area: toggle between your original text and the English translation to see exactly what the model will receive
  • The original prompt, detected language, and English translation are all preserved in the asset metadata
  • File names are generated from the translated English prompt (so "病院の建物" → hospital-building_opt1_var1.png)
  • Chat Studio passes prompts directly to the LLM (no translation) since models like Claude are natively multilingual
  • Type Studio text stays in your language (it's rendered on the image as-is)
  • All moderation pre-checks and content screening work on the translated English prompt for consistency

📌 1. What It Does

ArtSmoker works in two modes — standalone (no art style or theme setup needed, just describe and generate) and style-guided (upload your existing art, and every generation matches your visual identity). Both modes use the same studios and generation pipeline.

📝 Standalone Mode (Quick Start)

No style or theme setup needed — open the 2D Image Studio, Video Studio, or Type Studio and start creating immediately.

  1. Describe what you need — type a prompt like "hospital building" or "fire mage character", or use voice input. The AI decomposes your idea into visual components, enhances it with model-specific optimisations, and respects your creative intent through smart lock/vary controls. Write in any language — non-English prompts are auto-translated.
  2. Choose your models and settings — multi-select from all available text-to-image models (Amazon Bedrock + self-hosted on SageMaker), pick dimensions, quality tier, and region. Check multiple models for side-by-side comparison, or select one for focused generation. Cost estimate updates live.
  3. Get genuinely different options — the system generates up to 5 distinctly different creative concepts (varying outfit, mood, lighting, composition — not just camera angle), each with up to 5 seed variations (25 images total). User-specified details are locked; AI-inferred details are varied boldly. A visible Seed control makes batches reproducible — the same seed with the same final prompt and settings regenerates the same images, and clicking any result anchors to its seed so you can tweak one thing and branch from a favorite.
  4. Edit and refine — use inpainting, outpainting, erase, search & replace, or recolor directly in the Asset Viewer. Each edit creates a new version — the original is always preserved.
  5. Download game-ready files — PNG with transparent background + SVG, named descriptively (e.g. hospital-building_opt2_var3.png). Videos export as MP4.

📝 Style-Guided Mode (Match Your Art Style & Theme)

For teams that want every generated asset to match an existing art style — upload reference images and let the AI learn your visual identity first.

  1. Upload your game's art — import reference images from local directories (recursive scan, symlinked to avoid duplication) or S3 buckets (recursive listing with pagination). Smart deduplication runs automatically — removes rotation variants (barrel_N/E/S/W.png keeps only barrel_S.png) and animation frames (Idle0-Idle8 keeps only Idle). For example, a 747-file isometric asset pack deduplicates to ~99 unique objects. Supports: .png, .jpg, .jpeg, .gif, .bmp, .webp, .tiff, .tif, .tga, .ico, .svg, plus automatic texture extraction from 3D models (.glb, .gltf).
  2. AI learns your style — two-phase cohesion-aware analysis: first, a quick check determines whether your collection is unified, structurally consistent, or diverse. Then a deep analysis of the full reference set produces a metadata-rich style profile — colour palettes, line weights, lighting patterns, composition rules, and production conventions. If you provide generation hints, the AI receives them as "Artist's Guidance" so the analysis understands your intent, not just what's visible.
  3. Generate with style applied — when you select a style in the Image Studio, every prompt is automatically enhanced with your style's visual directives. A prompt like "hospital building" becomes a detailed generation instruction that includes your game's colour palette, perspective conventions, and rendering style.
  4. Everything from Standalone mode applies — multiple options, model comparison, editing, versioning, and game-ready downloads all work the same way, now guided by your art style.

Note

All generated content is produced by AI models and depends on the prompts and references you provide. Please review the Disclaimer regarding content quality, intellectual property, and applicable service terms before using generated assets in production.

📝 1.1 Features at a Glance

  • 🎨 Style Library — Upload art, AI learns your visual identity
  • 🖼️ 2D Image Studio — Generate images with options x variations, guided 3-step prompt workflow
  • 🎨 Prompt Designer — AI decomposes your prompt into editable visual components (subject, scene, lighting, colours) with lock/vary toggles per field, style integration, and smart asset type classification. Photorealistic, Character, Environment, and more
  • 🗂️ Collections (Set Generation) — Turn ONE prompt or a reference image into a coherent set of distinct assets ("a Viking-styled chess set", "20 JRPG characters"). Flip the Collection toggle in the Image Studio — from the Text tab (your prompt drives it) or the Image-Inspired tab (a reference image's look drives the art direction) — and a dedicated Collection Designer decomposes your brief into a shared art direction plus a roster of distinct, in-theme pieces, each with its own editable, model-agnostic prompt. Edit the art direction to recompose the whole set; lock, regenerate, add, or delete individual pieces; then generate them all (options × variations each) in one go, with a live design + projected-cost estimate. Pick one model for a cohesive set, or several to render every piece on each (a cross-model comparison). Optional hero- or reference-anchor cohesion style-locks the set to a chosen/supplied image. Collections appear as one card in the Gallery, opening a Collection viewer where each piece drills into the normal Asset Viewer (options / variations / versions, plus the piece's collection lineage); you can pin which version represents each piece, 3D the whole set at once, export it as a per-engine bundle, or load the whole collection back into the Image Studio to review and regenerate
  • 🎬 Video Studio — Text-to-video with model-specific prompt guidance (Nova Reel camera controls, Luma Ray natural language), multi-shot, image-to-video
  • ✍️ Type Studio — AI-designed text overlays with font picker
  • 💬 Chat Studio — Multi-model LLM chat with streaming, markdown, code highlighting, vision, sessions, context compaction
  • 📁 Unified Gallery — Masonry layout that shows each asset at its true aspect (portrait, square, landscape — never cropped). Browse images + videos, media filter (All / 2D Artwork / 3D Models / Video), search, full date-time-timezone stamps, download, delete. Assets that already have a generated 3D model carry a 3D badge, and the 3D Models filter surfaces just those
  • 📥 Import Image — Bring an existing image (any format) into the Gallery as a first-class asset. Auto-converted to PNG, tagged with an asset type you choose, and immediately editable and 3D-ready — everything (versioning, editing, image-to-3D) works exactly as it does for a generated image
  • ✏️ Image Editing — Inpainting, outpainting, erase, search & replace, recolor (in AssetViewer). Each mode has an AI Generate Prompt button: a vision model reads the image + its original prompt and proposes an edit prompt tailored to that mode and to the selected edit model (a descriptive caption for Stability editors, an instruction for Qwen-Image-Edit). Extend/Outpaint shows a live growing-frame preview with pixel rulers so you can see exactly how far the canvas will expand before committing. Instruction editors (Qwen-Image-Edit) support all five modes mask-free — including true canvas extension: ArtSmoker pre-pads the canvas, has the model complete only the new region, and blends your original pixels back untouched. Every edited version shows both model tags — the original generator and the editor that made that version
  • 📤 Export & Cutouts — Per-version export artefacts in the AssetViewer: a background-removed transparent PNG cutout plus true vector SVG traces (with and without background). Background removal is your choice per run: free on-device (rembg/u2net, no cloud cost) or the paid Amazon Bedrock remover — the same choice is offered when preparing images for 3D generation
  • 🔄 Real-time progress — SSE streaming with retry/throttle visibility
  • 🛡️ Smart moderation — Canary testing, auto model switching, AI-assisted rewriting
  • ⚙️ Model Registry — Admin UI organized by studio (Image, Video, Chat, Type, Shared), Bedrock discovery, custom model support
  • 📝 Prompt Templates — 28 editable LLM directive prompts, AI-assisted refinement, variable validation with auto-fix
  • 📦 Asset Versioning — Edit-in-place with version history (v1, v2, ...), version navigation, and per-version delete: remove just one version (others keep their numbers), with the viewer switching to the previous version — deleting the last version removes the whole asset
  • 💰 Cost Tracking — Estimated AWS spend per request, per session, per asset, computed from live per-region AWS pricing; self-hosted models show the GPU instance's hourly running rate + typical generation time (not a misleading per-image price)
  • 🌐 9-Language i18n — Full UI translation (EN, JA, ZH, KO, HI, RU, FR, ES, DE), auto-detect non-English prompts (English UI skips detection entirely), bilingual preview
  • 🔍 Custom Model Support — Discover fine-tuned, imported, and deployed custom Bedrock models automatically
  • 🔧 Self-Hosted Models — 1-Click Deploy — Browse a curated catalog of pre-tested open-source models (Qwen-Image, Qwen-Image-Edit, HunyuanImage 3.0, FLUX.2, FLUX.1, TripoSG, TRELLIS.2, and more), pick a GPU instance, and click Deploy. ArtSmoker handles everything: packaging the inference handler, configuring quantisation, selecting the right CUDA toolkit, setting up auto-scaling, registering CloudWatch alarms, and wiring async job tracking. Every model in the catalog has been validated end-to-end — from cold start through generation to gallery delivery — so you don't have to debug GPU drivers, memory overflows, or container compatibility. Supports BF16 + FlashInfer for best quality, NF4 for cost efficiency, multi-GPU auto-detection, auto-scales to zero ($0 idle), and the same model runs on different instance types without reconfiguration
  • 🧊 Image-to-3D Generation — Convert any Game Asset or Character image into a textured 3D mesh (GLB) with one click. Multi-view synthesis + texture baking produces game-engine-tailored assets. Interactive 3D viewer with orbit/zoom/pan
  • 🩹 Smart source completion for 3D — image-to-3D can only build what's visible, so a cropped character (legs cut off) becomes a legless mesh. Before generating, ArtSmoker vision-checks the source and, if it's cropped, offers to complete it via outpainting (an AI-suggested, fully editable prompt) — previews the before/after, re-reviews the result, lets you extend again or discard, and saves it as a new image version. Opt-in and non-blocking; well-framed images generate straight through
  • 🔄 Auto-Update — Version-gated on startup + a 24h periodic check; updates via git (checkout) or a tarball download-and-replace for installs without git, then restarts in place (supervised respawn / gunicorn reload) or offers a one-click Restart — never clobbers your data/ or .env (ARTSMOKER_AUTO_UPDATE=false to disable)

📝 1.2 Screenshots

2D Image Studio — Settings on the left with multi-select model dropdown, asset type, dimensions, and post-processing options. 3-step prompt workflow on the right with Prompt Designer and Generate Enhanced Prompt buttons. IP declaration and cost estimate at the bottom.

2D Image Studio — Settings, prompt workflow, and generation controls

2D Image Studio — Generation Results — Enhanced prompt shown above, multi-model comparison results below. Each model generates independently with per-model prompt optimization. Results display model name, dimensions, and generation cost.

2D Image Studio — Enhanced prompt and generation results

2D Image Studio — Model Comparison — Side-by-side comparison grid across all selected models (8 shown — Amazon Bedrock and self-hosted alike). Each option card carries its own variation filmstrip; the per-model negative prompt is shown for the selected option. Post-processing toggles (Remove Background, Convert to SVG, Upscale) apply to existing results without regenerating.

2D Image Studio — Multi-model comparison grid with variations

Image Inspiration (reference-guided) — Drop 1–3 reference images, say what you want, and pick how they're used: Match the reference (pixel-faithful edit on a deployed image-editing model) or Inspired by the reference (a vision AI writes the enhanced prompt — works with any model selection, options, and variations). The derived prompt is previewed and fully editable before generating.

Image Inspiration — reference images, instruction, and the editable enhanced-prompt preview

Image Inspiration — Results — The reference becomes a new creation (here, a caricature drawn from the reference photo), with the exact prompt sent to the model and the per-image cost recorded.

Image Inspiration — generated caricature results from a reference image

Prompt Designer — AI decomposes your prompt into editable visual components (Subject, Scene, Composition, Lighting, Style & Colors). Each field can be individually edited with lock/vary controls for genuinely distinct creative options.

Prompt Designer — Structured visual decomposition with editable fields

Prompt Designer — Color Palette — Named color palettes with hex swatches, style keywords, and quality level controls. The AI learns your visual identity and applies it consistently across all generations.

Prompt Designer — Color palette, style keywords, and quality controls

Style Library — Upload your game's existing art, AI analyzes the visual style and produces a metadata-rich prompt guide. Reference images are displayed with the full AI analysis and JSON style profile.

Style Library — AI style analysis with reference images

Style Library — Reference images, import options, and analysis data

Gallery — Unified view of all generated images and videos with media type filter, style filter, search, and sorting. Click any asset to open the full viewer. The Import Image button brings an existing image into the gallery — pick an asset type (Character/Game Asset enable 3D), and it's converted to PNG and made instantly editable and 3D-ready.

Gallery — Generated assets grid with filters

Asset Viewer — Full-size preview with tabbed interface (PNG, Edit, Export & Cutouts, Metadata, 3D Model), the image version bar, and direct PNG/SVG downloads. Zoom/fit/measure controls over the checkerboard-composited image.

Asset Viewer — Full-size preview with download options

Asset Viewer — Image Editing — Five edit modes: Fill/Replace, Remove, Extend, Find & Replace, Recolor. Shown: Extend with the measurement ruler, per-side pixel amounts, and the ✨ Generate Prompt button that reads the image and writes the edit prompt for you. Version history preserved — originals are never overwritten.

Asset Viewer — Extend editing with measurement ruler and AI-suggested prompt

Asset Viewer — Export & Cutouts — Per-version artefacts ready for your game, engine, or design tool: full-image vector SVG, background-removed cutout PNG, and cutout SVG. Background removal runs on-device for free (a paid Amazon Bedrock pass is optional).

Asset Viewer — Export & Cutouts with vector SVG and background-removed cutouts

After an outpainting round (v3 below), the same tab regenerates all three artefacts for the improved full-body version.

Asset Viewer — Export & Cutouts for the outpainted full-body version

Asset Viewer — Metadata — The full prompt lineage (your prompt → Prompt Designer decomposition → recomposed prompt → the model-tailored refined prompt), generation details, cost breakdown, and the complete version history.

Asset Viewer — Metadata with full prompt lineage and version history

The 3D pipeline's screenshots — generation, source review, game-engine-tailored export, and variants — are shown in section 1.9 (3D Model Generation) below, alongside the features they illustrate.

Video Studio — Settings on the left (model, generation mode, duration, region, cost estimate), prompt on the right. Supports Nova Reel (single shot, multi-shot auto/manual up to 2 minutes) and Luma AI Ray (aspect ratios, looping).

Video Studio — Settings and prompt

Video Studio — Generation in progress with AI-enhanced prompt

Video Studio — Completed video with thumbnail and recent videos

Video Player — Click a video to play it inline with full metadata (original prompt, AI-enhanced prompt, model, duration, region).

Video Player — Playing a generated video with metadata

📝 1.3 Two-Level Generation

For each prompt, the AI creates Options — fundamentally different design interpretations (e.g. for "a warrior": Viking berserker, Japanese samurai, tribal fighter, cyber-soldier, Greek hoplite). For each option, the image model produces Variations — different random seeds giving subtle visual differences. This gives artists a broad creative palette to choose from.

📝 1.4 Multi-Model Selection

The model dropdown supports checkbox-based multi-select — pick any combination of models for a single generation run:

  • Single model — check one model for focused generation (fastest, cheapest)
  • Multiple models — check 2-3 specific models for targeted comparison (e.g. SD 3.5 + FLUX.2 only)
  • All Available Models — toggle at the bottom selects/deselects all enabled models for a full side-by-side comparison

Each model runs independently: if stricter models block the prompt, you still get results from models that accepted it, with clear status labels (success, blocked by moderation, or failed) on each option card. The cost estimate updates live as you check/uncheck models.

An optional "Model-optimized prompts" toggle tailors the prompt to each model's strengths — prompts are rewritten per model (e.g. quality boosters for SD 3.5, natural language for FLUX.2, best-in-class text rendering cues for Qwen-Image).

📝 1.5 Reference-Guided Generation

Beyond writing a prompt from scratch, you can generate from 1–3 reference images plus an instruction — pick the mode with the segmented control at the top of the Image Studio prompt area:

  • Match the reference — keep the subject, product, or character from your reference and change the rest (theme, background, wardrobe, lighting) exactly as your instruction says. Ideal for consistent characters or product shots across scenes. This mode runs on a self-hosted instruction editor (Qwen-Image-Edit) and appears once it's deployed — if it isn't, ArtSmoker points you straight to deploy it from Custom Models (one click, same flow as the 3D pipelines). Commercial-safe (Apache-2.0).

  • Inspired by the reference — ArtSmoker's vision AI reads your reference(s) and instruction, writes the enhanced prompt(s), then generates with your normal text-to-image models — your full model selection, Options, and Variations are honored, with one distinct interpretation per option. The derived prompts are previewed as editable text before generating, so what you approve is exactly what runs. Always available — no deployment needed. Great for borrowing a look, palette, or composition without copying the subject.

  • Remix the reference — classic strength-based image-to-image: your reference's pixels go straight to a Bedrock model (Stable Diffusion 3.5 Large or Stable Image Ultra — registry capability-flagged) with a strength dial: subtle keeps the composition, palette, and mood nearly intact; wild treats it as loose inspiration. Set Options > 1 and they become a strength ladder — one card per strength, subtle to wild, side by side. Keeps layout, not identity (faces and products drift — use Match for identity-exact work). Output size follows the reference image. Always available — no deployment, no vision-analysis call.

All three modes require an instruction so you stay in control of what the reference is for. Reference-guided generation is separate from the Style Library (which analyzes many images into a reusable style profile) — use it for one-off, image-driven generations.

📝 1.6 Video Studio

Generate AI-powered videos and animations from text prompts. Supports Amazon Nova Reel (v1.0, v1.1) and Luma AI Ray (v2.0).

Feature Nova Reel Luma Ray v2
Max duration 120s (2 minutes) 9 seconds
Resolution 1280x720 720p / 540p
Aspect ratios 16:9 only 7 options (1:1, 16:9, 9:16, etc.)
Image-to-video Yes (start frame) Yes (start + end frame)
Looping video No Yes
Multi-shot control Yes (auto + manual) No
Price ~$0.08/sec ~$1.50/sec

How it works:

  1. Select a video model and configure duration, aspect ratio, region
  2. Enter a prompt — AI enhances it with cinematic vocabulary, camera movements, and temporal coherence cues
  3. Click Generate — the job runs asynchronously via StartAsyncInvoke, output goes to your configured S3 bucket
  4. Poll status every 5 seconds — on completion, thumbnail is extracted (via ffmpeg) and MP4 is downloaded locally (or streamed from S3)
  5. Videos appear in both the Video Studio's "Recent Videos" section and the unified Gallery

S3 bucket required: Video generation outputs to S3. You can configure via Video Settings in the UI (browse existing buckets or create new), or create one via CLI:

# Create an S3 bucket for video storage (replace REGION and YOUR_ORG)
aws s3api create-bucket --bucket artsmoker-video-YOUR_ORG --region us-east-1

# For regions other than us-east-1, add the LocationConstraint:
aws s3api create-bucket --bucket artsmoker-video-YOUR_ORG --region us-west-2 \
  --create-bucket-configuration LocationConstraint=us-west-2

Storage mode: download locally (default) or stream from S3 on demand.

Video prompt enhancement: The LLM adds camera movements (pan, zoom, dolly, tracking), lighting details, and temporal cues. Since video models don't support negative prompts, avoidance concepts are woven into the positive prompt naturally.

📝 1.7 Chat Studio

A full-featured LLM chat interface — like a self-hosted conversational AI, running on your own AWS account with no third-party data access.

80+ models from 16 providers — Claude (Sonnet, Opus, Haiku), Amazon Nova, Meta Llama, Mistral, Cohere, Qwen, DeepSeek, Google Gemma, NVIDIA Nemotron, and more. Plus any custom/imported models in your account. All discovered automatically via Sync from AWS.

Core features:

  • Streaming responses — real-time token-by-token rendering via Bedrock ConverseStream
  • Markdown rendering — headings, bold/italic, lists, tables, blockquotes, horizontal rules
  • Code blocks — syntax highlighting (highlight.js) with language badge + copy button
  • Per-message metrics — input/output tokens, latency, estimated cost, model used
  • Context window bar — visual fill indicator (green/amber/red) with used/max token count
  • Region switching — each model shows all available regions, pick the closest or cheapest

Session management:

  • Multiple concurrent sessions with auto-save
  • Inline rename, duplicate, delete, search/filter in sidebar
  • Export conversations as Markdown
  • Session totals: token count, estimated cost, message count

Advanced features:

  • System prompt templates — General Assistant, Coding Expert, Creative Writer, Game Designer, Data Analyst, Technical Writer
  • Vision/multimodal — drag-drop, file picker, or Ctrl+V paste images for vision-capable models
  • Context compaction — AI summarizes older messages to free context window space
  • Regenerate — re-run any AI response with the same prompt
  • Edit & resend — modify any user message and replay from that point
  • Fork — branch a conversation from any message into a new session

Pricing transparency: model picker shows cost per 1K tokens, pricing info bar shows estimated cost for 10K and 100K token conversations.

📝 1.8 Asset Type Awareness

The selected Asset Type fundamentally changes how the AI interprets your prompt — not just the image model, but every stage of the pipeline. When you type "hospital" and select different asset types, you get completely different outputs:

Type Composition Framing Technical Approach
Photorealistic Image (default) Natural, photograph-like framing — the subject in a contextually appropriate real setting. Real-world camera perspective: eye-level, shallow depth of field for portraits, wide for landscapes. Directed in photography language (golden hour, studio softbox, focal-length feel) with natural imperfections — skin texture, fabric wrinkles, weathering. Never illustration terms or render-engine jargon.
Game Asset Single isolated object on transparent background. No scene, no text, no UI. Straight-on or isometric, object fills 70-80% of frame. Clean sharp edges for bg removal, consistent top-left lighting, no ground shadows. Designed to compose with other game assets at various scales.
Character Full-body or 3/4-body figure, isolated on clean background. One character only. Character fills 60-75% vertical, head-to-toe, slightly off-center. Strong readable silhouette (identifiable from silhouette alone), expressive pose conveying personality, clear facial features and costume details.
Icon Single bold recognizable symbol, centered with generous padding. Maximum simplicity. Front-facing or slight 3/4 tilt, breathing room at edges. Must read clearly at 64x64 pixels. High contrast, 3-5 colors maximum, bold shapes, no thin lines or fine detail.
Marketing Banner Full scenic illustration with dramatic composition. Clean text-safe zone reserved on one side — no rendered text or typography. Wide cinematic feel, camera pulled back to show a scene. Rich saturated colors, dramatic lighting (rim light, volumetric rays), depth-of-field. The AI is explicitly instructed NOT to render text; the text-safe zone is left clean for post-production overlay in design tools (Figma, Canva, etc.).
Environment Full landscape with foreground/midground/background depth layers, leading lines. Wide establishing shot, horizon at upper or lower third. Atmospheric perspective (distant objects lighter/hazier), environmental storytelling through details, mood-setting lighting.

This matters at every stage:

  • "Preview Enhanced Prompt" button — When you click Compose, the AI uses the asset type to reshape your brief into a detailed generation prompt, combining your words with style guidelines and asset type directives. Your explicit intent always overrides style defaults. You can review the composed version before generating.
  • Concept generation — When generating multiple options, the AI creates N different design interpretations that all respect the asset type's structural rules. A Character option always has a readable silhouette; a Marketing Banner option always has a text-safe zone with no rendered text.
  • The result — Two images from the same prompt but different asset types will look nothing alike. A Game Asset "warrior" is a single centered character sprite. A Marketing Banner "warrior" is an epic battle scene with a clean zone for headline overlay.

📝 1.9 3D Model Generation (Image-to-3D)

Generate fully-textured 3D meshes from any 2D image — directly in the Asset Viewer. Select a Game Asset or Character image, open the 3D Model tab, and click Generate. The result is a game-engine-tailored GLB you can orbit, zoom, and download — with no manual modeling, UV unwrapping, or texture painting.

The end result, first: an ArtSmoker-generated character, exported as game-engine-tailored FBX and opened in stock Blender — the LOD chain (LOD0–LOD3) intact in the Outliner, textures bound, nothing re-rigged or fixed by hand. Everything below shows how to get here from a text prompt.

ArtSmoker FBX opened in Blender — LOD group hierarchy intact with textures bound

The generated model — orbit, inspect, download:

3D Model Generation — the generated soldier mesh viewed from multiple angles in the interactive 3D viewer

A single 2D character image (left, in the PNG tab) becomes a fully-textured 3D mesh you can rotate freely in-browser. The 3D Model tab now also lists the exact models & tools used to produce each asset (geometry model, texturing backend, output type, instance, and generation parameters) — persisted to the asset's metadata for full provenance.

Generating — the Asset Viewer's 3D Model tab: pick the deployed pipeline endpoint, quality tier (with est. time and cost), and advanced parameters. The licence panel shows each pipeline's terms, and Improve the Source vision-checks the image before you spend GPU time.

3D Model Generation — Settings and generation in the Asset Viewer

Improve the Source — Before generating, ArtSmoker measures the subject's silhouette and flags crops (here: cut off at the bottom edge), suggesting extension amounts and an AI-written outpaint prompt — extend, fill, or use the image as-is.

3D source review — automatic crop detection with suggested extension

Two pipelines — your choice. ArtSmoker offers two ways to turn an image into a textured 3D model. Deploy either (or both) from Custom Models; when both are live, you pick per generation in the Asset Viewer — each shows its est. cost, time, and licence so you decide informed:

Pipeline How it works Licence Commercial use Best for
TripoSG + texture backend TripoSG builds the mesh; a chosen texture backend (TRELLIS.2 / Hunyuan3D-Paint) paints it per backend (below) per backend Mixing geometry + a specific texturer
TRELLIS.2 (Full) One model generates both geometry and PBR texture (SLAT) MIT model + weights ⚠️ Non-commercial bake — the mesh export hard-imports nvdiffrast (NVIDIA 1-Way Commercial); review before commercial use Simplest path; prototypes & non-commercial work

Commercial use: the TripoSG + texture backend path is the commercial-safe route — its bake uses Kaolin (Apache-2.0). The full TRELLIS.2 pipeline is MIT for the model itself, but its o_voxel mesh bake hard-imports nvdiffrast (NVIDIA Source Code License, non-commercial for general users) and can't be swapped without patching upstream — so treat full-pipeline output as non-commercial unless you've cleared that dependency. The licence + every dependency is shown in the deploy dialog and again at generate time.

3D variants — Keep multiple 3D takes per image version (here TripoSG vs the TRELLIS.2 full pipeline), switch between them or set the default at any time; every variant records the exact models & tools that produced it.

3D variants — TripoSG and TRELLIS.2 takes side by side with full provenance

How the TripoSG pipeline works:

  1. Geometry extraction — a rectified flow transformer (TripoSG, 1.5B parameters, MIT-licensed) converts a single 2D image into a high-fidelity 3D mesh using a signed distance field (SDF) representation. Mesh density scales with the quality preset (up to ~1M faces at the highest octree resolution) for crisp detail on faces and gear.
  2. Texturing — the mesh is painted by a texture backend you choose at deploy time (default TRELLIS.2, Microsoft, MIT — a SLAT/voxel-conditioned texturer producing full PBR materials at a 4096² atlas).
  3. PBR output — exported as a GLB with embedded PBR maps, ready for physically-based rendering in any modern engine.

The TRELLIS.2 (Full) pipeline does the same end-to-end in a single model — no separate texturing step.

Licence in plain sight — at deploy AND at generate. Each deployable option shows its full licence and dependency breakdown in the deploy dialog — every model it pulls, that model's licence, and whether it's commercial-OK or gated — and you read and accept before deploying. At generate time, the Asset Viewer surfaces the licence again and confirms "accepted at deploy on <date>" (no second click needed):

Texture backend Licence Commercial use Best for
TRELLIS.2 (default) MIT ✅ Yes — requires a "Built with DINOv3" attribution in your product Production, commercial assets, highest quality
Hunyuan3D-Paint Tencent Community ❌ Non-commercial Research / non-commercial, exceptional faces

Background removal (the cutout step) uses BiRefNet (MIT) by default — fully commercial-clean — with a non-commercial alternative (RMBG) available as a disclosed opt-in. ArtSmoker never silently pulls a restricted dependency: anything gated or non-commercial is named, badged, and gated behind an explicit acceptance.

Output: Standard GLB with embedded PBR textures — imports directly into Unity, Unreal Engine, Blender, and other game engines. The interactive 3D viewer supports orbit, zoom, and pan for immediate inspection, and the 3D Model tab lists the exact models & tools used (geometry model, texturing backend, dependencies, instance, parameters) for full provenance.

Infrastructure: Both pipelines deploy via the same 1-click Custom Models flow, with the deploy-time picker showing each option's licence, dependency table, instance baseline, and est. cost/time. The full TRELLIS.2 pipeline's right-sized baseline is ml.g5.2xlarge (~$1.52/hr; a live run measured peak ~4.78 GB VRAM on the A10G's 24 GB and completed cleanly — chosen over the L40S g6e.xlarge, which is hugely over-provisioned for this workload at higher cost). Larger g6e sizes remain available as headroom upsells. Endpoints scale to zero when idle — $0 cost between jobs. First cold start builds the CUDA extensions once (then S3-cached for fast restarts). Before deploying a gated model, the dialog pre-checks HuggingFace access for every repo it pulls and shows a ✓/✗ per repo with the exact next step — so you never discover a missing licence acceptance minutes into a cold start.

Viewing the GLB: textures are encoded as WebP (EXT_texture_webp) to keep files compact — renders perfectly in the in-app viewer, Blender 4.x, three.js, and modern Unity/Unreal importers. macOS Preview/QuickLook doesn't support WebP-in-glTF and will show the model black; use the in-app viewer or any modern glTF tool.

Metric Value
Mesh quality Up to ~1M faces, full vertex normals
Texture resolution 4096² PBR atlas (base color + metallic-roughness + alpha)
Licensing Commercially-safe by default (TRELLIS.2 MIT + BiRefNet MIT); non-commercial backends offered with full disclosure
Supported asset types Game Asset, Character

📝 1.9.1 Engine-Ready Exports (GLB · FBX · USD)

3D Model viewer with per-variant tools and game-engine-tailored FBX/USD export options

Every generated 3D model can be exported prepared for your game engine, straight from the Asset Viewer's 3D tab:

  • Target engine — pick Generic (glTF, Y-up), Unreal Engine (Z-up), Unity, Godot, Maya, or 3ds Max. FBX and USD exports are oriented with the correct up/forward axes for that engine, so models import upright — no manual rotation fixes.
  • Optional prep, your choice (each an independent dropdown — nothing is forced):
    • Texture packing — per-engine texture sets: Unreal ORM (AO/Roughness/Metallic), Unity Metallic + Smoothness-in-alpha, Unity HDRP Mask Map. When selected, the export becomes a ZIP with the model plus a textures/ folder.
    • LODs — a decimated LOD0–LOD3 chain (100/50/20/5%) with a real FBX LOD group that Unreal auto-imports; the _LOD0…_LOD3 naming doubles as Unity's convention.
    • Collision — a convex hull or a CoACD convex decomposition, named per engine convention (UCX_* for Unreal auto-import, -convcolonly suffixes for Godot).
    • Lightmap UV2 — a second, smart-projected UV channel for baked lighting.
  • Two-step flow — buttons read Generate FBX/USD/GLB for a combination that doesn't exist yet; clicking converts server-side (a status line keeps you informed — large models can take a minute or two). Once built, the button flips to Download with a ✓ and delivers instantly. Every distinct combination is cached — nothing is ever regenerated.
  • Ready to download chips — the 3D tab lists every combination you've already generated for the current version, one click to re-download any of them.
  • The original GLB is sacred — "Download GLB (original)" always returns the untouched, byte-exact output of the generation pipeline. Processed exports (including a processed GLB with LODs/collision baked in) are separately-named files beside it.
  • Zero setup — conversions run server-side through a managed headless Blender: an existing installation is reused if present, otherwise a portable copy downloads automatically on first use (see the Model Settings → Maintenance tab for version and updates). End users install nothing.

📝 1.9.2 What to Expect from AI-Generated 3D — An Honest Guide

Image-to-3D is a young technology, and it's worth knowing what today's best models (including the ones ArtSmoker runs) genuinely deliver — and what they don't. The output is a dense, scanned-object-style mesh: up to ~1M unstructured triangles with baked PBR textures. Up close you'll notice a characteristic lumpy surface quality, and thin features (hair strands, straps, fabric fringes) are where AI geometry is weakest. There is no clean quad topology, no animation-friendly edge loops, and no rig — that's the state of the art across the industry, not a limitation unique to any one tool.

Where these assets shine — and where they need an artist:

Use case Ready to use?
Props, environment clutter, set dressing ✅ Yes — usable as-is
Background / mid-distance characters, crowds ✅ Yes — at distance the surface noise disappears; use the LOD chain
Prototyping, blockouts, previz, pitch demos ✅ Yes — arguably the strongest use case
Mobile / stylized games ✅ Often — the decimated LODs help
Hero characters, close-ups, animated characters ⚠️ A starting point — plan on artist retopology, cleanup, and rigging

What ArtSmoker adds on top of the raw mesh is that everything arrives correctly packaged for your engine — right up-axis per target, LOD chain, collision proxies, engine-specific texture packing — so the remaining work is creative, not plumbing.

Inspecting exports in Blender (or another DCC tool)? Two things will look strange — both are correct:

  • Generated with LODs? The file contains 4 stacked copies of the model (LOD0–3). Viewed together they shimmer (z-fighting) and look noisy — hide LOD1–3 in the Outliner and judge quality from LOD0 alone. Game engines show exactly one LOD at a time, so this never happens in-engine.
  • Generated with collision? A white, blocky shell of UCX_* meshes encloses the model — that's the physics proxy, not your asset. Hide those objects to see the textured model inside. Engines import them as invisible collision automatically.

📝 1.9.3 Using a Model Commercially — Who to Pay, and How

If you love a model's output and want to use it commercially, the path depends on how its creator monetizes. ArtSmoker shows every model's license at deploy time; this is the "what do I do next?" companion. (Verified against vendor sites and HuggingFace license files in August 2026 — licensing changes fast, so always confirm the vendor's current terms. Informational only — see the Disclaimer.)

The four patterns you'll meet:

  1. Already yours (Apache-2.0 / MIT) — commercial use is included, free. There is no license product to buy; the creator monetizes via their own hosted API instead. Your only obligation is notice/attribution compliance.
  2. Free until you're big (community licenses) — commercial use is included below a threshold (revenue or monthly active users). Above it, the license itself instructs you to request an enterprise grant from the vendor — a sales conversation, not a store.
  3. Buy the license, keep the weights — the HuggingFace weights are non-commercial, but the creator sells a separate self-host commercial license. Once you hold it, the same weights you already deployed become legal for commercial use — nothing changes technically in ArtSmoker.
  4. The gate is the paywall — the HuggingFace repo itself is gated; a commercial agreement with the vendor unlocks your HF account's access. ArtSmoker's HuggingFace-token + gated-repo pre-check flow then works as-is.
Creator / models Pattern Your next step for commercial self-hosting
Alibaba — Qwen-Image, Qwen-Image-Edit 1 Nothing to buy (Apache-2.0). Retain the license notices.
Microsoft — TRELLIS.2 · VAST — TripoSG 1 Nothing to buy (MIT). Note: upstream deps (e.g. Meta DINOv3) carry their own gates/terms.
Black Forest Labs — FLUX.2 [klein] 4B 1 Nothing to buy — Apache-2.0, free commercial use.
Stability AI — SD 3.5 (self-hosted) 2 Free incl. commercial under $1M total annual revenue (the HF gate acceptance is the license). Above it the license auto-terminates — request an Enterprise license at stability.ai/enterprise. "Powered by Stability AI" attribution is mandatory at every tier.
Tencent — HunyuanImage 3.0 / Hunyuan3D 2 Free incl. commercial below the threshold — note it's per-model: HunyuanImage 3.0 = 100M MAU; Hunyuan3D-2.1 = only 1M MAU (above: email hunyuan3d@tencent.com). No grant at all in the EU, UK, or South Korea, regardless of size.
Black Forest Labs — FLUX.1/.2 [dev], Kontext 3 Buy a FLUX Commercial Weights License (self-serve at dashboard.bfl.ai/licensing; tiers are image-volume-capped subscriptions). You keep using the same HF weights. Mind the obligations: usage reporting, output filtering, no exposing the model as an API/resale; new model versions are not automatically covered outside Enterprise.
Bria — FIBO, RMBG-2.0 4 The HF gate grants non-commercial access immediately; commercial self-hosting requires a paid agreement with Bria (purchase form linked from each model card / bria.ai). No free-commercial threshold exists. Once granted, deploy through ArtSmoker exactly as before.

How this plays with ArtSmoker: procurement almost never changes anything technical. For patterns 1–3 the weights you deploy are identical before and after — what changes is the contract you hold (keep your license record; ArtSmoker's deploy dialog records your acceptance of the weights license, but commercial grants bind you directly with the vendor). For pattern 4, once the vendor approves your HuggingFace account, ArtSmoker's existing gated-repo access check turns green and deployment proceeds normally. When a vendor releases a new model version, re-check whether your grant covers it (Stability's does automatically; BFL's generally doesn't outside Enterprise; Tencent issues a fresh license text per version).

📌 2. Prerequisites

  • Python 3.11+ (3.12, 3.13, 3.14 all work)
  • AWS CLI configured with working credentials
  • IAM permissions for Bedrock access (see below)

📝 2.1 AWS Credentials

ArtSmoker uses boto3's standard credential resolution, so any of the following methods work:

Method Best for How
Environment variables CI/CD, containers AWS_ACCESS_KEY_ID + AWS_SECRET_ACCESS_KEY
Shared credentials file Local development ~/.aws/credentials via aws configure
Named profile Multiple accounts Set ARTSMOKER_AWS_PROFILE=myprofile or AWS_PROFILE
AWS SSO Enterprise SSO aws configure sso
IAM Instance Profile EC2, ECS, App Runner Attach an IAM role to the instance — no credentials needed on the machine
ECS Task Role ECS/Fargate containers Assign a task execution role with the required permissions

Quick check that credentials are working:

aws sts get-caller-identity

Note

On EC2 and other AWS compute services, you don't need to configure explicit credentials. Attach an IAM Instance Profile with the required permissions, and boto3 picks it up automatically via the instance metadata service.

📝 2.1.1 Verify Bedrock Access

Confirming credentials work (sts:GetCallerIdentity) only verifies identity — it does not confirm you have Bedrock permissions. ArtSmoker uses multiple Bedrock APIs, so a quick listing test alone is not sufficient. The most reliable check:

# Test 1: Can you list models? (requires bedrock:ListFoundationModels)
aws bedrock list-foundation-models --region us-east-1 --query "modelSummaries[0].modelId" --output text

# Test 2: Can you invoke an image model? (requires bedrock:InvokeModel)
aws bedrock-runtime invoke-model --region us-west-2 \
  --model-id stability.sd3-5-large-v1:0 \
  --content-type application/json --accept application/json \
  --body '{"prompt":"test","aspect_ratio":"1:1"}' \
  /dev/null 2>&1 && echo "InvokeModel: OK" || echo "InvokeModel: FAILED"

# Test 3: Can you use the Converse API? (authorizes via bedrock:InvokeModel)
# (Substitute any Claude model ID you have access to — e.g. the current Sonnet
#  inference profile from Test 1's list; the exact version rolls over time.)
aws bedrock-runtime converse --region us-west-2 \
  --model-id us.anthropic.claude-sonnet-5 \
  --messages '[{"role":"user","content":[{"text":"hi"}]}]' \
  --inference-config '{"maxTokens":1}' \
  --query "output.message.content[0].text" --output text 2>&1 && echo "Converse: OK" || echo "Converse: FAILED"

# Test 4: Can you list custom models? (requires bedrock:ListCustomModels)
aws bedrock list-custom-models --region us-east-1 \
  --query "modelSummaries[0].modelName" --output text 2>&1 && echo "ListCustomModels: OK" || echo "ListCustomModels: no custom models (or permission denied)"

If Tests 1-3 pass, your core permissions are set. Test 4 is needed only for custom model discovery. If Test 1 passes but Tests 2-3 fail, your IAM policy allows listing but not invoking — update it using the permissions table below.

📝 2.2 IAM Permissions

Your IAM user, role, or instance profile needs these permissions:

Permission Used for
bedrock:InvokeModel Image generation, editing, post-processing — and all non-streaming LLM calls (the Converse API authorizes via this action; there is no separate bedrock:Converse action)
bedrock:InvokeModelWithResponseStream Streaming LLM responses (Chat Studio) — the ConverseStream API authorizes via this action
bedrock:InvokeModelWithBidirectionalStream Voice transcription (optional — app works without it)
bedrock:StartAsyncInvoke Video generation (async invocation; the operation itself authorizes via bedrock:InvokeModel, which is also granted)
bedrock:GetAsyncInvoke Poll video generation job status
bedrock:ListAsyncInvokes List video generation jobs
bedrock:ListFoundationModels Foundation model discovery (Sync from AWS)
bedrock:ListCustomModels Discover fine-tuned custom models in your account
bedrock:ListImportedModels Discover imported models in your account
bedrock:GetCustomModel Read custom model details (base model, status)
bedrock:GetImportedModel Read imported model details (architecture, status)
bedrock:ListProvisionedModelThroughputs Find invocable custom models with provisioned throughput
bedrock:ListCustomModelDeployments Find custom models with on-demand deployments
bedrock:CreateInference (or policy AmazonBedrockMantleInferenceAccess) Amazon Bedrock Mantle — frontier models reachable only via the Mantle endpoint (OpenAI GPT‑5.x, Claude Mythos, GLM, Grok, Qwen, Gemma…). Missing it affects only those models; Claude via Converse keeps working.
account:ListRegions Scan only your account's enabled regions during Sync (fast, no errors on opt‑in regions). Optional — falls back to scanning all regions.
account:GetRegionOptStatus Read per‑region opt‑in status (companion to account:ListRegions). Optional.
s3:CreateBucket Create S3 bucket for video storage (optional, via UI)
s3:PutObject / s3:GetObject / s3:DeleteObject / s3:ListBucket Video output storage and retrieval; async-job inputs/outputs and 3D artefacts for Custom Models
s3:ListAllMyBuckets The S3 bucket browser in Video Settings / Style import (optional)
s3:GetLifecycleConfiguration / s3:PutLifecycleConfiguration Auto-expiry rules for async-job temp files on the deployment bucket (Custom Models)
aws-marketplace:Subscribe Auto-subscription on first use of third-party models (incl. third-party Mantle models)
aws-marketplace:ViewSubscriptions Check existing model subscriptions
sts:GetCallerIdentity Startup credential validation; also underpins the locally-signed Mantle bearer token
pricing:GetProducts Fetch model pricing during Sync from AWS (optional)
sagemaker:* Self-hosted custom models on Amazon SageMaker (optional — only if using Custom Models)
application-autoscaling:* (register/deregister targets, put/delete/describe policies) Scale-to-zero + scale-from-zero for Custom Models endpoints (optional — only for Custom Models)
cloudwatch:PutMetricAlarm / DeleteAlarms / DescribeAlarms The scale-from-zero backlog alarm for Custom Models endpoints (optional)
logs:DescribeLogStreams / FilterLogEvents / GetLogEvents / PutRetentionPolicy Endpoint readiness detection + log retention for Custom Models (optional)
servicequotas:GetServiceQuota / RequestServiceQuotaIncrease GPU instance quota check / one-click increase request in the deploy dialog (optional)
ecr:DescribeRepositories Resolve the SageMaker Deep Learning Container image at deploy (optional)
iam:CreateServiceLinkedRole (scoped to sagemaker.application-autoscaling.amazonaws.com) First-ever endpoint auto-scaling in the account auto-creates AWSServiceRoleForApplicationAutoScaling_SageMakerEndpoint (optional — skip if that role already exists)
iam:PassRole Allow Amazon SageMaker to use your role (optional — only for Custom Models)
iam:CreateRole / iam:AttachRolePolicy / iam:PutRolePolicy Auto-create Amazon SageMaker execution role on first deploy (optional — only for Custom Models)
iam:GetRole / iam:UpdateAssumeRolePolicy Auto-configure existing role for Amazon SageMaker trust (optional)
secretsmanager:CreateSecret / UpdateSecret / GetSecretValue / DescribeSecret / DeleteSecret Encrypted storage for HuggingFace tokens on gated models (optional — auto-cleaned on teardown)

Quickest setup (managed policies — broadest access):

# Option A: Attach managed policies to your IAM user (simplest for local development)
aws iam attach-user-policy --user-name YOUR_USERNAME \
  --policy-arn arn:aws:iam::aws:policy/AmazonBedrockFullAccess

# Amazon Bedrock Mantle endpoint — needed for frontier models (OpenAI GPT-5.x,
# Claude Mythos, GLM, Grok, etc.). Skip only if you won't use Mantle-only models.
aws iam attach-user-policy --user-name YOUR_USERNAME \
  --policy-arn arn:aws:iam::aws:policy/AmazonBedrockMantleInferenceAccess

# Add S3 access for video storage
aws iam attach-user-policy --user-name YOUR_USERNAME \
  --policy-arn arn:aws:iam::aws:policy/AmazonS3FullAccess

Scoped setup (tighter permissions — recommended for production):

# Create a scoped IAM policy with only the permissions ArtSmoker needs
aws iam create-policy --policy-name ArtSmokerAccess --policy-document '{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "Bedrock",
      "Effect": "Allow",
      "Action": [
        "bedrock:InvokeModel",
        "bedrock:InvokeModelWithResponseStream",
        "bedrock:InvokeModelWithBidirectionalStream",
        "bedrock:StartAsyncInvoke",
        "bedrock:GetAsyncInvoke",
        "bedrock:ListAsyncInvokes",
        "bedrock:ListFoundationModels",
        "bedrock:ListCustomModels",
        "bedrock:ListImportedModels",
        "bedrock:GetCustomModel",
        "bedrock:GetImportedModel",
        "bedrock:ListProvisionedModelThroughputs",
        "bedrock:ListCustomModelDeployments"
      ],
      "Resource": "*"
    },
    {
      "Sid": "BedrockMantleInference",
      "Effect": "Allow",
      "Action": ["bedrock:CreateInference", "bedrock:GetInference", "bedrock:ListInferences", "bedrock:GetInferenceProfile"],
      "Resource": "*"
    },
    {
      "Sid": "EnabledRegions",
      "Effect": "Allow",
      "Action": ["account:ListRegions", "account:GetRegionOptStatus"],
      "Resource": "*"
    },
    {
      "Sid": "S3VideoStorage",
      "Effect": "Allow",
      "Action": ["s3:CreateBucket", "s3:PutObject", "s3:GetObject", "s3:ListBucket", "s3:DeleteObject", "s3:HeadBucket", "s3:GetLifecycleConfiguration", "s3:PutLifecycleConfiguration"],
      "Resource": ["arn:aws:s3:::artsmoker-*", "arn:aws:s3:::artsmoker-*/*"]
    },
    {
      "Sid": "Marketplace",
      "Effect": "Allow",
      "Action": ["aws-marketplace:Subscribe", "aws-marketplace:ViewSubscriptions"],
      "Resource": "*"
    },
    {
      "Sid": "Utility",
      "Effect": "Allow",
      "Action": ["sts:GetCallerIdentity", "pricing:GetProducts", "s3:ListAllMyBuckets"],
      "Resource": "*"
    },
    {
      "Sid": "SageMakerCustomModels",
      "Effect": "Allow",
      "Action": [
        "sagemaker:CreateModel", "sagemaker:CreateEndpointConfig", "sagemaker:CreateEndpoint",
        "sagemaker:UpdateEndpoint", "sagemaker:DeleteModel", "sagemaker:DeleteEndpointConfig",
        "sagemaker:DeleteEndpoint", "sagemaker:DescribeEndpoint", "sagemaker:DescribeEndpointConfig",
        "sagemaker:InvokeEndpoint", "sagemaker:InvokeEndpointAsync"
      ],
      "Resource": "arn:aws:sagemaker:*:*:*artsmoker*"
    },
    {
      "Sid": "SageMakerList",
      "Effect": "Allow",
      "Action": ["sagemaker:ListModels", "sagemaker:ListEndpointConfigs"],
      "Resource": "*"
    },
    {
      "Sid": "CustomModelsRuntime",
      "Effect": "Allow",
      "Action": [
        "application-autoscaling:RegisterScalableTarget", "application-autoscaling:DeregisterScalableTarget",
        "application-autoscaling:DescribeScalableTargets", "application-autoscaling:PutScalingPolicy",
        "application-autoscaling:DeleteScalingPolicy", "application-autoscaling:DescribeScalingPolicies",
        "cloudwatch:PutMetricAlarm", "cloudwatch:DeleteAlarms", "cloudwatch:DescribeAlarms",
        "logs:DescribeLogStreams", "logs:FilterLogEvents", "logs:GetLogEvents", "logs:PutRetentionPolicy",
        "servicequotas:GetServiceQuota", "servicequotas:RequestServiceQuotaIncrease",
        "ecr:DescribeRepositories"
      ],
      "Resource": "*"
    },
    {
      "Sid": "AutoScalingServiceLinkedRole",
      "Effect": "Allow",
      "Action": "iam:CreateServiceLinkedRole",
      "Resource": "arn:aws:iam::*:role/aws-service-role/sagemaker.application-autoscaling.amazonaws.com/AWSServiceRoleForApplicationAutoScaling_SageMakerEndpoint",
      "Condition": {"StringLike": {"iam:AWSServiceName": "sagemaker.application-autoscaling.amazonaws.com"}}
    },
    {
      "Sid": "SageMakerRoleManagement",
      "Effect": "Allow",
      "Action": ["iam:CreateRole", "iam:AttachRolePolicy", "iam:PutRolePolicy", "iam:GetRole", "iam:UpdateAssumeRolePolicy", "iam:PassRole"],
      "Resource": ["arn:aws:iam::*:role/ArtSmoker*"]
    },
    {
      "Sid": "SecretsManagerHFTokens",
      "Effect": "Allow",
      "Action": ["secretsmanager:CreateSecret", "secretsmanager:UpdateSecret", "secretsmanager:GetSecretValue", "secretsmanager:DescribeSecret", "secretsmanager:DeleteSecret"],
      "Resource": "arn:aws:secretsmanager:*:*:secret:artsmoker/*"
    }
  ]
}'

# Attach to your IAM user (replace YOUR_ACCOUNT_ID and YOUR_USERNAME)
aws iam attach-user-policy --user-name YOUR_USERNAME \
  --policy-arn arn:aws:iam::YOUR_ACCOUNT_ID:policy/ArtSmokerAccess

Note

Two things to adjust for your account: (1) The S3 statement is scoped to buckets named artsmoker-* — if you point ArtSmoker at a bucket with a different name (Video Settings lets you pick any existing bucket), widen that Resource to your bucket's ARN. (2) SageMaker resources ArtSmoker creates are named artsmoker-*, so the scoped Resource works as-is; the sagemaker:List* actions don't support resource scoping and live in their own statement.

Missing something at runtime? ArtSmoker monitors every AWS call: if one is denied, it logs the exact service:Operation that failed and raises a persistent in-app notice naming the action to add — so a permission gap never fails silently.

Tip

For EC2/ECS/App Runner — create an IAM role instead of attaching to a user. See the EC2 Deployment section for the complete role creation commands. No access keys needed — boto3 auto-discovers the role from the instance metadata service.

Note

Bedrock models are available by default in all commercial AWS regions — no manual enablement step is needed. On first invocation of a third-party model (Anthropic, Stability AI), AWS automatically initiates a marketplace subscription in the background (requires the aws-marketplace permissions above). Anthropic models require a one-time First Time Use form completion.

📝 2.3 Optional: SVG Conversion Tools

SVG conversion uses external CLI tools (not Python packages). Without them, SVG output falls back to a Pillow-based raster-in-SVG wrapper — functional but not true vector output.

Tool Purpose macOS Linux (Debian/Ubuntu) Windows
vtracer Primary SVG (color vector tracing) pip install vtracer or cargo install vtracer pip install vtracer or cargo install vtracer pip install vtracer or cargo install vtracer or pre-built binaries
potrace Fallback SVG (monochrome tracing) brew install potrace sudo apt install potrace Download from potrace.sourceforge.net

Verify installation:

# Check SVG conversion tools
which vtracer && echo "vtracer: OK" || echo "vtracer: not installed (optional)"
which potrace && echo "potrace: OK" || echo "potrace: not installed (optional)"

📝 2.4 Optional: Video Thumbnail & Metadata Tools

Video Studio generates MP4 videos via Amazon Nova Reel and Luma AI Ray. To extract thumbnails (first frame as JPEG) and video metadata (duration, resolution, FPS), ffmpeg and ffprobe must be installed on the machine running the ArtSmoker backend.

Without ffmpeg:

  • Videos still generate and play correctly (streamed from S3 or downloaded as MP4)
  • Thumbnails will be missing — the Gallery and Video Studio show a black placeholder instead of a preview image
  • Video metadata (duration, resolution) won't be displayed
Tool Purpose macOS Linux (Debian/Ubuntu) Windows
ffmpeg Thumbnail extraction + video metadata brew install ffmpeg sudo apt install ffmpeg Download from ffmpeg.org/download or winget install ffmpeg

Note

ffprobe is included with ffmpeg — no separate install needed. ArtSmoker checks for ffmpeg at runtime and falls back gracefully if it's not found — video generation works either way, you just won't get thumbnails.

Verify installation:

ffmpeg -version 2>&1 | head -1 && echo "ffmpeg: OK" || echo "ffmpeg: not installed (optional)"
ffprobe -version 2>&1 | head -1 && echo "ffprobe: OK" || echo "ffprobe: not installed (optional)"

📌 3. Installation

📝 3.1 macOS

git clone https://github.com/niravdd/ArtSmoker.git && cd ArtSmoker

# Option A: With virtual environment (recommended)
python3 -m venv .venv
source .venv/bin/activate
pip install -r backend/requirements.txt

# Option B: Without virtual environment (system-wide install)
pip3 install -r backend/requirements.txt

Note

On macOS, python3 and pip3 are available via Homebrew (brew install python) or the Xcode command-line tools. If you see "command not found", install Python from python.org or via brew install python@3.12.

📝 3.2 Linux (Debian/Ubuntu)

# Install Python if needed
sudo apt update && sudo apt install python3 python3-pip python3-venv

git clone https://github.com/niravdd/ArtSmoker.git && cd ArtSmoker

# Option A: With virtual environment (recommended)
python3 -m venv .venv
source .venv/bin/activate
pip install -r backend/requirements.txt

# Option B: Without virtual environment
pip3 install --user -r backend/requirements.txt

Note

On some Linux distros, pip install outside a venv requires the --user flag or --break-system-packages (PEP 668). Using a venv avoids this entirely.

📝 3.3 Windows

git clone https://github.com/niravdd/ArtSmoker.git
cd ArtSmoker

# Option A: With virtual environment (recommended)
python -m venv .venv
.venv\Scripts\activate
pip install -r backend\requirements.txt

# Option B: Without virtual environment
pip install -r backend\requirements.txt

Note

On Windows, use python (not python3). Install Python from python.org — check "Add to PATH" during installation. If PowerShell refuses to run .venv\Scripts\activate ("running scripts is disabled"), allow it once with Set-ExecutionPolicy -Scope CurrentUser RemoteSigned, or use .venv\Scripts\activate.bat from cmd. The Type Studio font picker detects fonts from C:\Windows\Fonts (system font detection is currently macOS/Linux only — Windows users can use global or style-specific custom fonts).

📌 4. Running

📝 4.1 Solo Development (All Platforms)

Single-process with auto-reload on file changes — ideal for one developer working locally:

# With venv (activate first)
source .venv/bin/activate          # macOS / Linux
.venv\Scripts\activate             # Windows

uvicorn backend.main:app --reload
# Without venv (if installed system-wide)
uvicorn backend.main:app --reload

# Or if uvicorn isn't on PATH
python3 -m uvicorn backend.main:app --reload     # macOS / Linux
python -m uvicorn backend.main:app --reload       # Windows

Open http://localhost:8000 — the frontend is served by FastAPI, no separate web server needed.

On startup, the console shows AWS credential validation results. If something's wrong, you'll see a clear error box. You can also check http://localhost:8000/api/health for the status.

Logs. In addition to the console, ArtSmoker writes a full, append-only log to logs/artsmoker.log by default, so you can review a past session after the app has closed. Each run is framed by a session banner (launch time, version, pid, host) and closed with a shutdown banner (stop time, duration). To change the path or turn it off:

ARTSMOKER_LOG_FILE=/var/log/artsmoker/app.log uvicorn backend.main:app   # custom path
ARTSMOKER_LOG_TO_FILE=false uvicorn backend.main:app                      # disable file logging

(Or set log_to_file / log_file in a local .env. With multiple workers, every worker appends to the same file.)

Auto-restart (optional, all platforms). The commands above all work as-is. To also let ArtSmoker restart itself in place — so an auto-update, or the in-app Restart button, reloads new code without you re-launching — start it under the built-in cross-platform supervisor instead:

python -m backend.main            # add --host / --port as needed; Ctrl-C stops it cleanly

This works on every OS, including Windows. The supervisor runs the app in a child process and respawns it on a restart request. (Running under gunicorn or a service manager like systemd gives the same in-place restart — see §4.2 and §6.)

📝 4.2 Multi-User / Shared Test Box / Production (macOS / Linux)

For any environment with more than one concurrent user — whether a shared dev/test box, staging, or production — use gunicorn with multiple workers:

# Install gunicorn (one-time, in addition to requirements.txt)
pip install gunicorn

# Run with gunicorn (multi-worker, handles concurrent users)
gunicorn backend.main:app \
  -w 2 \
  -k uvicorn.workers.UvicornWorker \
  --bind 0.0.0.0:8000 \
  --timeout 300
Flag Purpose
-w 2 2 worker processes (increase for heavier load)
-k uvicorn.workers.UvicornWorker Use uvicorn's async worker class
--bind 0.0.0.0:8000 Listen on all interfaces (not just localhost)
--timeout 300 5-minute timeout for large batch generations with retries

Tip

gunicorn is Linux/macOS only. On Windows, use uvicorn backend.main:app --host 0.0.0.0 --port 8000 --workers 2 for multi-worker serving.

Note

Safe for concurrent users. All server writes — image/version metadata and the model & prompt registries — are written atomically and serialized across worker processes (POSIX file locks), so simultaneous edits from multiple collaborators on a shared box never corrupt a file or lose an update. File logging works the same across workers — each appends to the one logs/artsmoker.log.

📝 4.3 EC2 / Cloud Deployment

Recommended: t3.small (~$15/month) for 1-2 concurrent users.

Step 1: Create an IAM role for the EC2 instance (run from your local machine):

# Create the IAM role with EC2 trust policy
aws iam create-role --role-name ArtSmokerEC2Role \
  --assume-role-policy-document '{
    "Version": "2012-10-17",
    "Statement": [{
      "Effect": "Allow",
      "Principal": {"Service": "ec2.amazonaws.com"},
      "Action": "sts:AssumeRole"
    }]
  }'

# Attach the ArtSmoker policy (use the scoped policy from section 2.2, or the managed policy)
aws iam attach-role-policy --role-name ArtSmokerEC2Role \
  --policy-arn arn:aws:iam::aws:policy/AmazonBedrockFullAccess
aws iam attach-role-policy --role-name ArtSmokerEC2Role \
  --policy-arn arn:aws:iam::aws:policy/AmazonS3FullAccess

# Create an instance profile and attach the role
aws iam create-instance-profile --instance-profile-name ArtSmokerEC2Profile
aws iam add-role-to-instance-profile \
  --instance-profile-name ArtSmokerEC2Profile \
  --role-name ArtSmokerEC2Role

Step 2: Launch an EC2 instance (or attach the profile to an existing one):

# Attach to an existing running instance
aws ec2 associate-iam-instance-profile \
  --instance-id i-YOUR_INSTANCE_ID \
  --iam-instance-profile Name=ArtSmokerEC2Profile

Step 3: Install and run on the instance (SSH into the instance):

# Install (one-time)
sudo yum install -y python3 python3-pip git   # Amazon Linux
# sudo apt install -y python3 python3-pip python3-venv git   # Ubuntu

git clone https://github.com/niravdd/ArtSmoker.git && cd ArtSmoker
python3 -m venv .venv
source .venv/bin/activate
pip install -r backend/requirements.txt
pip install gunicorn

# Optional: install ffmpeg for video thumbnails
sudo yum install -y ffmpeg   # Amazon Linux
# sudo apt install -y ffmpeg   # Ubuntu

Step 4: Run as a systemd service (persistent, auto-restarts):

# Create the service file
sudo tee /etc/systemd/system/artsmoker.service > /dev/null << 'EOF'
[Unit]
Description=ArtSmoker
After=network.target

[Service]
WorkingDirectory=/home/ec2-user/ArtSmoker
ExecStart=/home/ec2-user/ArtSmoker/.venv/bin/gunicorn backend.main:app \
  -w 2 -k uvicorn.workers.UvicornWorker --bind 0.0.0.0:8000 --timeout 300
Restart=always
User=ec2-user

[Install]
WantedBy=multi-user.target
EOF

# Enable and start
sudo systemctl daemon-reload
sudo systemctl enable artsmoker
sudo systemctl start artsmoker

# Verify it's running
sudo systemctl status artsmoker

# View logs
sudo journalctl -u artsmoker -f

Open http://YOUR_INSTANCE_IP:8000 — ensure your EC2 security group allows inbound TCP 8000.

📝 4.4 First Steps After Setup

After ArtSmoker is running, complete these steps to get the best results:

1. Sync models from AWS — Open Model Settings (gear icon in any studio) → click Sync from AWS. This discovers all available image, video, and chat models across all Bedrock regions. Takes 30-60 seconds. Only needed once, or when AWS adds new models.

2. Review and customize prompt templates — This is the most impactful configuration you can do. Open Model Settings → Prompt Templates tab. ArtSmoker uses 28 editable directive prompts that control how the AI behaves:

Template What it controls
Image Prompt Refinement How your text descriptions are turned into detailed image generation prompts
Multi-Concept Generation How multiple creative options are generated from a single idea
Style Analysis How reference images are analyzed to learn your art style
Content Moderation How strict the pre-check and rewrite system is
Video Enhancement How video prompts are enriched with camera movements and lighting
Text Layout How Type Studio designs text positioning on images

Each template can be:

  • Edited directly — modify the instructions to match your team's needs
  • Enhanced with AI — select any LLM model, optionally add instructions (e.g., "optimize for pixel art"), and click "Enhance with AI". Review the suggestion, then Accept or Dismiss
  • Reset to default — restore the original at any time

Templates are organized by studio (Image Studio, Style Library, Content Safety, Video Studio, Type Studio, Chat Studio, Translation) with friendly descriptions of what each one controls.

Variable safety: Templates use {curly_brace} variables (e.g., {user_prompt}, {model_name}) that are substituted at runtime. If you accidentally remove a required variable, ArtSmoker will:

  1. Block the save and show which variables are missing
  2. Offer "Fix & Save" — an LLM automatically inserts the missing variables back into your edited text in the right places
  3. Verify the fix before saving

Templates load from backend/prompt_templates.json — the runtime source of truth. Your edits are saved to backend/prompt_templates.user.json (gitignored) and overlaid on top, so an update or git pull never overwrites your customizations. If the JSON is missing or corrupted, or a new template ships in code, it self-heals: the built-in code seed regenerates/backfills only the missing entries, never overwriting existing ones.

Tip

Start by reviewing the Image Prompt Refinement and Creative Options templates. These have the biggest impact on output quality. If your team specializes in a particular art style (e.g., pixel art, watercolor, isometric), add those preferences directly into the templates so every generation benefits.

3. Set up a style profile (optional) — Go to Style Library, create a new style, upload reference images, and click Analyze. This teaches ArtSmoker your visual identity.

4. Choose your language — Click a language button in the nav bar (EN | 日 | 中 | 한 | हिं | РУ | FR | ES | DE) if you prefer a non-English interface.

📌 5. Architecture

┌─────────────────────────────────────────────┐
│  Browser (SPA)                              │
│  Vanilla JS + Tailwind CSS                  │
└──────────────────────┬──────────────────────┘
                       │ HTTP / SSE
                       ▼
┌─────────────────────────────────────────────┐
│  FastAPI Backend (Python)                   │
│                                             │
│  /api/styles      Style CRUD + import       │
│  /api/generate    Two-level generation      │
│  /api/type-studio Text overlay + fonts      │
│  /api/video       Video generation + jobs   │
│  /api/chat        LLM chat + sessions       │
│  /api/gallery     Asset browsing + export   │
│  /api/browse      File/S3 browser           │
│  /api/admin       Model registry + templates│
│  /api/refine-prompt  Prompt + translate      │
│  /api/transcribe  Voice-to-text             │
└────────────┬────────────────────┬───────────┘
             │                    │
             ▼                    ▼
┌──────────────────────┐  ┌──────────────────────────┐
│  us-west-2           │  │  us-east-1               │
│                      │  │                          │
│  Claude Sonnet       │  │  Nova Sonic (voice)      │
│  Claude Opus         │  │  Nova Reel (video)       │
│  SD 3.5 Large        │  │                          │
│  Stable Image Ultra  │  │                          │
│  Stable Image Core   │  │                          │
│  Stability AI (post) │  │                          │
└──────────────────────┘  └──────────────────────────┘ ... (other regions)
             │
             ▼
┌──────────────────────┐
│  Local Storage        │
│  data/styles/         │
│  data/images/      │
│  data/video/          │
│  data/chat/           │
└──────────────────────┘

📌 6. Usage

📝 6.1 Workflow Overview

                            ┌─────────────────┐
                            │   ArtSmoker     │
                            └────────┬────────┘
                                     │
       ┌───────────┼───────────┼───────────┼───────────┐
       │           │           │           │           │
       ▼           ▼           ▼           ▼           ▼
  ┌──────────┐ ┌────────┐ ┌────────┐ ┌──────────┐ ┌────────┐
  │  Style   │ │  2D    │ │ Video  │ │   Type   │ │  Chat  │
  │ Library  │ │ Image  │ │ Studio │ │  Studio  │ │ Studio │
  │          │ │ Studio │ │        │ │          │ │        │
  │ Upload   │ │Generate│ │Generate│ │ Add text │ │ Multi- │
  │ Analyze  │ │ images │ │ videos │ │ to imgs  │ │ model  │
  │ Set fonts│ │        │ │        │ │          │ │ LLM    │
  │          │ │        │ │        │ │          │ │ chat   │
  └────┬─────┘ └───┬────┘ └───┬────┘ └────┬─────┘ └────────┘
             │              │            │               │
             │    ┌─────────┴────────────┴─────────┐     │
             │    │  Style selected? (optional)    │     │
             └───►│  Enhances output               │◄────┘
                  └─────────┬──────────────────────┘
                                    │
                                    ▼
                          ┌─────────────────┐
                          │    Gallery      │
                          │                 │
                          │ Browse all      │
                          │ Search/filter   │
                          │ Select & delete │
                          └────────┬────────┘
                                   │
                    ┌──────────────┼──────────────┐
                    │              │              │
                    ▼              ▼              ▼
            ┌──────────────┐ ┌──────────┐ ┌──────────────┐
            │ Download     │ │ Reload   │ │ Add Text     │
            │ PNG / SVG    │ │ in 2D    │ │ in Type      │
            │              │ │ Image    │ │ Studio       │
            │              │ │ Studio   │ │              │
            │              │ │ (refine &│ │ (overlay     │
            │              │ │  regen)  │ │  text)       │
            └──────────────┘ └──────────┘ └──────────────┘

Three entry points, one unified gallery:

  • Start with a style — upload reference art in the Style Library, let AI analyze it, then generate in any studio. The style guides all output.
  • Start without a style — jump straight into 2D Image Studio, Video Studio, or Type Studio. AI uses its best judgement.
  • Start from the Gallery — pick any previously generated asset and reload it in the appropriate studio for refinement, add text to it, play a video, or download as PNG/SVG/MP4.

All generated assets (images, videos, text overlays, standalone text) land in the unified Gallery. Nothing is overwritten — each generation creates new assets.

📝 6.2 Generation Pipeline

User prompt: "hospital building"
         │
         ▼
┌────────────────────────────────────────────────────────┐
│ 1. Prompt Composition            Claude Sonnet (1 opt) │
│    (optional "Compose" button)   or Opus (2-5 options) │
│    + style + asset type                                │
└────────────────────────┬───────────────────────────────┘
                         │
                         ▼
┌────────────────────────────────────────────────────────┐
│ 2. Canary Test                                         │
│    Single image tests moderation                       │
│    Pass? ──► Full batch    Fail? ──► Model switch      │
│                                  or rewrite suggestion │
└────────────────────────┬───────────────────────────────┘
                         │
                         ▼
┌────────────────────────────────────────────────────────┐
│ 3. Parallel Image Generation                           │
│    Up to 5 options × 5 variations = 25 images          │
│    ThreadPool (3-5 workers)                            │
│    Retry with exponential backoff (3 attempts)         │
│    SSE progress streaming to browser                   │
│    Cooperative cancellation on moderation block        │
└────────────────────────┬───────────────────────────────┘
                         │
                         ▼
┌────────────────────────────────────────────────────────┐
│ 4. Post-Processing (per image, optional)               │
│    Remove Background ──► Stability AI ($0.07/img)      │
│    Upscale ──► Stability AI Creative Upscale ($0.60)   │
│    SVG ──► vtracer / potrace / Pillow (free, local)    │
└────────────────────────┬───────────────────────────────┘
                         │
                         ▼
┌────────────────────────────────────────────────────────┐
│ 5. Storage                                             │
│    data/images/{asset_id}/                          │
│    ├── asset.png (transparent background)              │
│    ├── asset.svg (optional)                            │
│    └── metadata.json (full prompt lineage)             │
│    Smart filenames: prompt-slug_opt1_var2.png          │
└────────────────────────────────────────────────────────┘

📝 6.3 Content Moderation Flow

User clicks Generate
         │
         ▼
┌──────────────────────┐
│ Pre-Check enabled?   │
│ (Prompt Pre-Check    │
│  toggle, on by       │
│  default)            │
└───┬────────────┬─────┘
  Yes            No
    │            │
    ▼            │
┌──────────┐     │
│ Claude   │     │
│ Sonnet   │     │
│ screens  │     │
│ prompt   │     │
└───┬────┬─┘     │
 Issues? No      │
    │    └──────►│
    ▼            │
┌──────────────┐  │
│ Indigo       │  │
│ dialog:      │  │
│ • Switch     │  │
│ • Rewrite    │  │
│ • Proceed    │  │
│ • Cancel     │  │
└──┬───────────┘  │
   │◄────────────┘
   ▼
┌──────────────────────┐
│ Canary test          │
│ (1 image to model)   │
└───┬────────────┬─────┘
 Blocked        Pass
    │            │
    ▼            ▼
┌──────────┐  ┌──────────┐
│ Try alt  │  │ Full     │
│ models   │  │ batch    │
└───┬────┬─┘  │ runs     │
 Works?  No   └──────────┘
    │    │
    ▼    ▼
Emerald    Amber
dialog     dialog
(switch    (rewrite →
 or         enhanced
 rewrite)   prompt area)

📝 6.4 2D Image Studio (Generate Assets)

The 2D Image Studio uses a guided 3-step workflow:

Step 1 — Describe your idea: Type a prompt in the textarea. The placeholder shows a realistic example that changes based on your selected Asset Type (e.g., "A young female warrior in ornate silver armor..." for Character, or "A misty Japanese garden at dawn..." for Environment). Use voice input (mic button) to dictate instead of typing.

Step 2 — Prompt Designer (optional): Click 🎨 Prompt Designer to decompose your prompt into structured visual components. The AI analyzes your prompt and breaks it into editable sections:

  • Subject — character description, clothing, accessories, pose, expression
  • Scene — setting, background, props, time of day
  • Composition — camera angle, framing, depth of field
  • Lighting — key light, fill/rim light, mood
  • Style & Colors — art style, quality level, and a named color palette with hex swatches

Each field can be individually edited. Generate Enhanced Prompt recomposes your edits into a flat recomposed prompt (shown read-only in Step 2) and then automatically generates the Enhanced AI Prompt for Step 3.

Before the Prompt Designer opens, an AI asset type classification runs — if your prompt describes a scene but you selected "Game Asset", a dialog suggests switching to "Environment" or "Character". This ensures the Prompt Designer decomposes with the correct context.

Step 3 — Enhanced prompt preview (optional): Click Generate Enhanced Prompt to see the model-optimized prompt before generating. The AI takes the recomposed prompt from Step 2 and enhances it with model-specific guidance (anatomy, materials, lighting, prompt structure). You can edit the enhanced prompt before generating. If you used the Prompt Designer in Step 2, this is auto-populated.

Prompt pipeline: User Prompt → Decompose → Recompose (recomposed_prompt) → Enhance with model guidance (enhanced_prompt) → Image Model. For multiple options, the enhancement step generates N distinct interpretations from the same recomposed base. All three levels are stored in metadata.

Generate: Click Generate at any point — Steps 2 and 3 are optional. If you skip them, Generate auto-decomposes, recomposes, and enhances your prompt before proceeding. Prompt Pre-Check (on by default) screens the prompt for moderation issues before generation.

Additional controls:

  • Asset Type — select in the sidebar. Changes the prompt placeholder and affects how the AI interprets your prompt. The system suggests switching if it detects a mismatch.
  • Art Style — select a style profile to guide generation with your visual identity.
  • Dimensions, Options, Variations — configure output size and how many creative concepts to generate.
  • Post-Processing — Remove Background, Upscale, SVG conversion (applied after generation).
  • IP Declaration — assert ownership or licensing for strict model compatibility.
  • Model Settings — view/edit model configuration, discover available Amazon Bedrock models.

Generation progress is streamed in real time via SSE — the UI shows which image is being generated (e.g. "Generating images... 12/25"), elapsed time, and current pipeline stage. If the API is throttled, you'll see "API throttled — waiting to retry..." with the delay, then "Retrying... (attempt 2/3)" — each image retries up to 3 times with exponential backoff so large batches don't lose variants to transient throttling.

Generated results survive navigation — switching tabs and back preserves the 2D Image Studio's DOM state. Only the reset button clears it.

Smart content moderation: When your prompt is blocked by a model's content moderation filters, ArtSmoker handles it progressively through three colour-coded dialogs:

  • Indigo (Pre-Check) — before generation, an AI pre-screens your prompt against the selected model's known sensitivity. If issues are detected, you see the specific concerns and can: switch to a recommended model, rewrite the prompt for the current model, proceed anyway, or cancel.
  • Emerald (Model Switch) — after a generation block, if an alternative model accepts your prompt as-is, ArtSmoker shows which model works and why. One-click to switch. Full attempt log available ("View N model tests").
  • Amber (Rewrite) — when all models reject, an AI-generated rewrite is offered in an editable textarea with specific issues listed. A verified/unverified badge indicates whether the rewrite passed canary testing.

Prompt rewrite behaviour: In all three dialogs, choosing "Rewrite" never overwrites your original prompt. The rewritten version appears in the enhanced prompt area below your original text, with a persistent amber disclaimer: "This rewrite is an attempt to make the prompt compatible — it is still subject to the model's own moderation assessment and may be rejected." You review and edit the enhanced prompt, then click Generate when satisfied. Your original prompt is always preserved in history and metadata.

Common triggers include copyrighted IP names and character references, violence/weapon language, and adult content references. Tip: the "Preview Enhanced Prompt" button often produces prompts that pass moderation naturally, since the AI rephrases in descriptive terms.

Smart canary testing: Before generating the full batch, ArtSmoker sends a single "canary" image request to test the prompt against the model's moderation filters. If the canary is blocked, the batch stops immediately (1 wasted API call instead of N×M×3). If the canary passes, remaining tasks run in parallel with cooperative cancellation — if any task hits a moderation block, the rest skip their API calls automatically.

📝 6.5 Use a Style Profile

  1. Go to the Style Library tab.
  2. Click Create New Style — enter a name and optionally add generation hints. In the create modal, use the "Import References From" section with Local and S3 browse buttons to select a source directory or bucket path. Browsing opens a server-side file/directory browser modal (single-click selects an item, double-click navigates into directories). Imported references are auto-analyzed on creation.
  3. Local directory imports scan recursively and symlink images (no duplication, portable relative links); 3D models (.glb/.gltf) get their embedded textures auto-extracted. S3 imports list recursively and download locally. Up to 100 reference images per style (supported formats are listed in section 1's Style-Guided Mode).
  4. Two-phase cohesion-aware analysis: a quick cohesion check (is the collection unified, structurally consistent, or diverse?) guides a deep vision analysis of a representative reference subset — extracting palette, materials, proportions, edge treatment, shadow/lighting, and detail-level attributes, plus ~200-word generation hints. Your own generation hints are fed in as "Artist's Guidance", so the profile captures intent, not just what's visible. (The full analysis mechanics — phases, models, subset selection — are in SPEC.md.)
  5. In the style detail view, use "Import & Analyze" to add more references and trigger analysis in one step. Drag-and-drop upload is also supported and auto re-analyzes when new images are added.
  6. "Re-Analyze Style" appears after the initial analysis, letting you manually re-run analysis at any time.
  7. Generation hints are part of the analysis context — the AI receives both reference images and your hints as "Artist's Guidance" when analyzing, so the style profile understands intent, not just visual appearance. Editing generation hints also triggers automatic re-analysis.
  8. Back in the 2D Image Studio, select your style from the dropdown — all generated assets will match its visual identity (palette, perspective, rendering style, mood).

📝 6.6 Style Analysis Flow

┌──────────────────────────────────────────┐
│ Create / Import style                    │
│ (reference images uploaded or imported)  │
└────────────────────┬─────────────────────┘
                     │
                     ▼
┌──────────────────────────────────────────┐
│ Phase 1: Cohesion Check                  │
│ Claude Sonnet — 8 images — ~$0.01        │
│ Determines: high / medium / low          │
│   high   = unified style                 │
│   medium = shared structure, diff themes │
│   low    = diverse collection            │
└────────────────────┬─────────────────────┘
                     │
                     ▼
┌──────────────────────────────────────────┐
│ Phase 2: Full Analysis                   │
│ Claude Opus — up to 20 images            │
│ Guided by cohesion level                 │
│ + Artist's Guidance (user hints)         │
│ Extracts 9 style attributes              │
└────────────────────┬─────────────────────┘
                     │
                     ▼
┌──────────────────────────────────────────┐
│ Phase 3: Hint Generation                 │
│ Claude Sonnet — 200-word hints           │
│ 8 dimensions: perspective, rendering,    │
│ materials, palette, proportions, edges,  │
│ shadow/lighting, detail level            │
└────────────────────┬─────────────────────┘
                     │
                     ▼
┌──────────────────────────────────────────┐
│ Stored in profile.json                   │
│ ~$0.14 total per style analysis          │
│ Used in all future generation            │
└──────────────────────────────────────────┘

📝 6.7 Type Studio

Add text to images or generate standalone text assets with AI-designed typography.

  • Two modes: "On Image" composites text onto a gallery image; "Standalone" renders text on a transparent background.
  • Multi-line text editor with per-line font selection, positioning controls, and voice input (mic button per line — dictate text via Nova Sonic transcription).
  • AI-designed layouts — the AI suggests colors, sizes, positions, and effects (shadow, outline, glow). Request 1–5 layout options for different creative directions. The LLM model used for layout is configurable (Complex LLM for best quality, Fast LLM for cheaper) — reads from the registry categories.
  • Font picker with live preview — style fonts, 8 bundled fonts (Roboto, Open Sans, Lato, Montserrat, Playfair Display, Oswald, Raleway, Source Code Pro), system fonts, and client-side detected fonts (via Local Font Access API or canvas probing).
  • Pre-Processing / Post-Processing — same workflow as 2D Image Studio, with an "Apply" button for post-processing. SVG conversion is on by default.
  • Click to zoom — clicking the result preview opens the AssetViewer with full zoom/pan, metadata, download, and image editing tools.
  • Results are saved as new gallery assets (originals are never overwritten).

📝 6.8 Gallery

  • Unified view of all generated images and videos in a masonry layout (each asset shown at its true aspect — portrait, square, or landscape — never centre-cropped), with a Media filter (All / 2D Artwork / 3D Models / Video). The 3D Models filter shows only assets that already have a generated 3D model, and those assets carry a 3D badge on their tile.
  • Search bar for instant filtering across all assets (prompts, styles, models).
  • Multi-select with checkboxes for bulk delete (handles both image and video assets). Deletions are batch-aware — surviving siblings track how many variants were removed, so reloading a partial batch in the Image Studio shows "X of Y images remaining (Z deleted)".
  • Assets load immediately with an in-memory metadata cache. Sorted newest-first.
  • Pagination support (limit/offset) for large collections.
  • Gallery auto-refreshes when you navigate back to it, and after any edit or video generation completes.
  • Video cards display a thumbnail with a play overlay, VIDEO badge, and duration indicator. Click to open the video player modal.
  • Contextual action buttons per asset based on type: "2D Studio" (indigo) to reload in the image studio, "Add Text" (emerald) to open in Type Studio, "Edit in Type Studio" (purple) for text assets.
  • Click any image to open the AssetViewer modal with:
    • Zoom/pan — mouse wheel to zoom, drag to pan, Fit/1:1 buttons with active mode highlighting.
    • Edit tab — inpaint, erase, outpaint, search & replace, or recolor the image directly. Two kinds of editor are offered per mode: mask-based (Stability) — paint a mask with the brush tool, enter a prompt, and apply; and mask-free instruction editors (Qwen-Image-Edit, when deployed) — just describe the change in words, no mask needed. The brush controls hide automatically for a mask-free model. Choose the editing model, apply; default replaces the original image, uncheck "Replace original" to save as a new asset (every edit preserves version history).
    • Previous / Next — arrow buttons and keyboard left/right to navigate through the list without closing the viewer.
    • Full metadata: original prompt, AI-improved prompt, generation prompt, negative prompt, style, asset type, image model (friendly names), dimensions, seed, batch ID, option/variation index, IP declaration status, filename, and creation date.
  • Style snapshot: Each asset stores a snapshot of the style used at generation time (name, description, hints, analysis). If the original style is later deleted, the asset retains the full context. Backward compatible — older assets without snapshots display normally.

📝 6.9 Voice Input

Click the microphone button next to the prompt editor to dictate your prompt. The audio is sent to Nova Sonic for transcription.

Note

Voice transcription requires Nova Sonic's bidirectional streaming API, which depends on a compatible boto3 version and model access enabled in us-east-1. If the streaming API is not available, the service returns a placeholder acknowledgment. Full real-time transcription works when Nova Sonic streaming is properly configured.

📝 6.10 View State Preservation

Navigation order: Style Library → 2D Image Studio → Type Studio → Video Studio → Gallery. Switching between views preserves each view's DOM state. Generated results, form inputs, and scroll positions survive navigation. The amber reset button in 2D Image Studio and Video Studio is the only way to clear their state.

📝 6.11 Model Management

All AI model configuration is centralized in backend/model_registry.json — the single source of truth. Models, regions, pricing, quality tiers, and format templates are all stored here and managed through the UI or API:

  • Click "Model Settings" in the sidebar of any studio to open the admin modal — it opens to the relevant tab for that studio.
  • 9 tabs organized by studio:
    • Image Studio — Image generation models (SD 3.5 Large, Stable Image Ultra, Stable Image Core, plus self-hosted FLUX, HunyuanImage, Qwen-Image), regions, quality tiers, prompt limits, moderation strictness
    • Video Studio — Video models (Nova Reel, Luma Ray), S3 bucket settings, regions, pricing
    • Chat Studio — Discovered chat/LLM models (80+ from 16 providers), context windows, vision capability, pricing per 1K tokens
    • Type Studio — LLM model for text layout generation (Complex or Fast LLM)
    • Shared Studio — Cross-studio LLM categories (Fast LLM, Complex LLM, Fallback LLM, Voice), post-processing models (Remove Background, Upscale)
    • Custom Models — the self-hosted model catalog: deploy, monitor, and tear down SageMaker endpoints (see section 6.12)
    • Prompt Templates — 28 editable LLM directive prompts organized into 6 workflow sections (see section 4.4)
    • Registry JSON — Raw JSON editor for the full model registry
    • Maintenance — managed tools status (e.g. the headless Blender used for FBX/USD export: path, version, on-demand update)
  • All sections are collapsible with Show All / Hide All toggles for quick navigation.
  • LLM categories and post-processing use dropdown model pickers (populated from discovered models) — not raw text fields.
  • Sync from AWS: Scans all Bedrock-supported AWS regions (discovered dynamically), auto-registers new image, video, and chat models, updates regional availability, fetches per-model pricing from the AWS Pricing API, and disables models no longer available. A live progress overlay streams each region as it's scanned. This is the only action that calls AWS discovery APIs — all other operations read from the cached registry.
  • Always on the newest Claude: Each Sync automatically rolls your Fast LLM to the newest Claude Sonnet and Complex LLM to the newest Claude Opus available in your account, so you're never stranded on a deprecated model — no manual config needed. If you manually pick a specific model for a category, it's pinned and the auto-roll leaves it alone (it just notifies you when a newer one appears).
  • Custom model discovery: Sync also discovers fine-tuned custom models (ListCustomModels), imported models (ListImportedModels), and models with on-demand deployments (ListCustomModelDeployments) or provisioned throughput (ListProvisionedModelThroughputs). Custom models inherit their format family from the base model automatically.
  • Auto-discovery: New foundation models are registered with enabled=true — the admin can disable them. Existing models get their available_regions and Bedrock metadata (modalities, lifecycle, ARN) updated automatically.
  • Styled confirmation dialogs: All destructive actions (Sync, delete, reset) use custom styled modals — no browser confirm() popups.
  • Changes are persisted immediately to model_registry.json via the Admin API.
  • The registry is backward compatible — existing assets reference model keys (e.g. sd35_large), not raw Bedrock model IDs.

📝 6.12 Self-Hosted Models (Custom Models on Amazon SageMaker)

ArtSmoker can deploy open-source AI models on Amazon SageMaker in your own AWS account, extending your capabilities beyond what Amazon Bedrock offers. These run alongside Bedrock models and appear in the same studio dropdowns.

Extensible model catalog: Ships with a built-in catalog of open-source models spanning image generation, upscaling, background removal, depth estimation, segmentation, and video. Adding a new model requires only a catalog entry — no code changes. You can also add custom models via the UI (+ Add Model). The catalog and available models evolve over time.

Deployment options:

  • Async (scale-to-zero) — pay only when generating. Scales to zero when idle ($0 cost), scales up automatically on new request. Cold start ~5-10 min.
  • Always-On — instant responses, ~$1.41/hr (ml.g5.xlarge)

How to deploy: Model Settings → Custom Models tab → click Deploy. The SageMaker container pulls model weights directly from HuggingFace at startup — no multi-GB local download required.

CPU offloading: Large diffusion models use intelligent CPU offloading to fit on smaller GPU instances. Each model's catalog entry specifies the strategy — model_cpu_offload (keeps active layers on GPU) or sequential_cpu_offload (aggressive per-layer offload for very large models). Applied automatically by the inference handler.

Async generation with Pending Jobs: Self-hosted models generate asynchronously. A Pending Jobs panel appears in the 2D Image Studio showing active jobs with progress indicators. Completed images arrive in the Gallery automatically — no polling or page refresh needed.

HuggingFace token management: Gated models require a Read-only HuggingFace token. The token is stored encrypted in AWS Secrets Manager in your account, managed via the UI (set/update/delete), and shared across all models that need it. Tokens are automatically cleaned up when you tear down all models.

Gated-access pre-check: Before a gated deploy, the dialog probes every HuggingFace repo the model pulls (its own weights plus any dependencies) using your stored token, and shows a ✓/✗ per repo with the exact next step — accept this repo's licence on HuggingFace, or add a token. Deploy stays blocked until every required repo is reachable, so a forgotten licence acceptance fails fast in the dialog instead of minutes into a cold start.

Setup: Add Amazon SageMaker and Secrets Manager permissions to the same IAM role you already use for Bedrock — no separate role or environment variable needed. ArtSmoker auto-discovers your role on EC2/ECS, or auto-creates an ArtSmokerSageMakerRole if needed.

# Add Amazon SageMaker permissions to your existing ArtSmoker role (one command)
aws iam attach-role-policy --role-name ArtSmokerEC2Role \
  --policy-arn arn:aws:iam::aws:policy/AmazonSageMakerFullAccess

Python dependency: huggingface_hub>=0.23 (install with pip install huggingface_hub)

📝 6.13 Image & Video Generation Models

All models are discovered dynamically from the registry — not hardcoded. The Image Studio dropdown is populated from GET /api/admin/models/image-options and the Video Studio dropdown from GET /api/admin/models/video-options on page load. Any model registered and enabled in the registry appears automatically.

The Image Model dropdown is the primary selection. Below it, a smart summary line shows the active region, quality tier, and per-image cost. An expandable Advanced section lets you override:

  • Quality — models that support quality tiers (a Standard/Premium price split) show a dropdown; models without tiers show "Default". Tiers are declared per model in the registry via quality_options.
  • Region — shows regions where the selected model is available, sorted cheapest-first with pricing. "Auto" selects the cheapest region.

A cost estimate updates dynamically based on all selections (model × quality × region × options × variations).

Format families: Models are invoked through a generic invoker that reads request templates from the registry (format_families) — generation, editing, post-processing, and video are all template-driven. Adding a new Bedrock image model requires zero code changes: register it (via auto-discovery or the admin API) with the correct format family. The full family catalog lives in SPEC.md.

Model-optimized prompt engineering: Prompts are automatically structured as descriptive captions (not commands) following AWS documentation. Negation words are removed from the main prompt and exclusion terms are sent as a separate negative prompt. The prompt is truncated to each model's specific prompt_limit from the registry.

Note

Moderation sensitivity varies by model and is tracked in the registry (moderation_strictness). The Amazon Bedrock Stability models (SD 3.5 Large, Stable Image Ultra, Stable Image Core) apply AWS platform moderation and are tuned "moderate"; self-hosted models (FLUX, HunyuanImage, Qwen-Image) run in your own account with no platform-imposed content filter. ArtSmoker handles blocks automatically — when a prompt is rejected, the system tries alternative models ordered by strictness before suggesting a rewrite.

📌 7. Tech Stack

Layer Technology
Backend FastAPI (Python 3.11+), boto3, Pydantic
Frontend Vanilla JS, Tailwind CSS (CDN)
AI (LLM) Claude Sonnet (fast tasks), Claude Opus (complex tasks)
AI (Image) Stable Diffusion 3.5 Large, Stable Image Ultra, Stable Image Core (Amazon Bedrock); FLUX.2/FLUX.1, HunyuanImage 3.0, Qwen-Image (self-hosted on SageMaker)
AI (Post-processing) Stability AI (Remove Background, Creative Upscale)
AI (Chat) 80+ LLMs from 16 providers via Bedrock ConverseStream (Claude, Nova, Llama, Mistral, etc.)
AI (Video) Nova Reel v1.0/v1.1 (up to 2min), Luma AI Ray v2 (up to 9s)
AI (Voice) Nova Sonic (speech-to-text via bidirectional streaming)
i18n Custom t() function, ~1,500 keys × 9 languages, reverse-lookup DOM translation
SVG Conversion vtracer (primary), potrace (fallback), Pillow (last resort)
Text Rendering Pillow (shadow, outline, glow effects)
Storage Local filesystem (S3-ready interface)
Dev No-cache middleware for static files; client-side error logging via POST /api/log

No build step required for the frontend.

📌 8. Security Model

ArtSmoker is designed as a local/trusted-network development tool — it runs on the developer's own machine or a private EC2 instance. The security model reflects this:

  • No authentication — all API endpoints are open. Appropriate for local development and private team deployments.
  • Filesystem browser — the GET /api/browse/local endpoint allows browsing any directory the server process can access. This is intentional for importing reference art from your machine.
  • Font serving — path traversal protection validates that font file requests stay within expected directories.
  • S3 access — S3 browsing and imports use the server's AWS credentials. The user can access any S3 bucket their IAM role permits.

Warning

Do not expose ArtSmoker to untrusted networks without adding authentication and path restrictions. See the Deployment Roadmap in SPEC.md for production hardening guidance (Phase 4 adds Cognito authentication).

📌 9. API

Interactive docs at http://localhost:8000/docs (Swagger UI).

Key endpoints:

Endpoint Purpose
Generation
POST /api/generate/ Generate assets (options × variations) with SSE streaming
POST /api/generate/post-process Apply processing to existing assets
POST /api/generate/edit Image editing: inpaint, outpaint, erase, search-replace, etc. Accepts source image, mask, prompt, model.
POST /api/generate/suggest-edit-prompt AI "Generate Prompt" for the Edit tab: reads the image + original prompt and returns an edit prompt for a given mode, styled for the target edit model (caption vs. instruction)
POST /api/generate/analyze-moderation Analyze a moderation-blocked prompt and suggest a safe rewrite
Styles
POST /api/styles/ Create a style profile
POST /api/styles/{id}/import Bulk-import references from a local folder or S3 URI
POST /api/styles/{id}/analyze Trigger AI style analysis
Prompt
POST /api/refine-prompt/ Preview a refined prompt
POST /api/transcribe/ Voice-to-text (Nova Sonic)
Gallery
GET /api/gallery/ Browse generated assets (supports limit/offset pagination)
GET /api/gallery/batch/{batch_id} Reconstruct full options × variations structure for a batch
DELETE /api/gallery/ Bulk delete assets
Type Studio
POST /api/type-studio/preview Render text overlay preview
POST /api/type-studio/suggest AI layout suggestion for text
GET /api/type-studio/fonts List available fonts
Browse
GET /api/browse/local?path=~ Browse local directory contents
GET /api/browse/s3/buckets List available S3 buckets
GET /api/browse/s3?bucket=name&prefix=path Browse S3 bucket contents
Chat
POST /api/chat/stream Stream LLM response via SSE (Bedrock ConverseStream)
GET /api/chat/models List all available chat models (foundation + custom + imported)
POST /api/chat/sessions Create a new chat session
GET /api/chat/sessions List chat sessions
GET /api/chat/sessions/{id} Load a full session (messages + metadata)
PUT /api/chat/sessions/{id} Update session (title, messages, model, temperature)
DELETE /api/chat/sessions/{id} Delete a session
POST /api/chat/sessions/{id}/duplicate Duplicate a session
GET /api/chat/sessions/{id}/export Export session as Markdown
GET /api/chat/sessions/{id}/search?q= Search within a session's messages
POST /api/chat/compact Compact older messages via LLM summarization
POST /api/chat/generate-title Auto-generate a session title from first exchange
Video
POST /api/video/generate Start async video generation job
GET /api/video/status/{job_id} Poll video generation job status
GET /api/video/jobs List all video generation jobs
GET /api/video/{id}/mp4 Serve video MP4 file
GET /api/video/{id}/thumbnail Serve video thumbnail
DELETE /api/video/{id} Delete a video
Admin
GET /api/admin/models Get full model registry (LLMs, image models, post-processing)
GET /api/admin/models/image-options Enabled text-to-image models for the dropdown (with pricing, quality tiers, regions). Accepts ?region= filter.
GET /api/admin/regions Cached list of Bedrock-supported AWS regions (no AWS calls)
PATCH /api/admin/models/category/{name} Update an LLM category config
PATCH /api/admin/models/image/{key} Update an image model config
POST /api/admin/models/image Add a new image model
POST /api/admin/discover/refresh-all Full refresh: discover regions + scan models + fetch pricing + prune stale data. The ONLY endpoint that calls AWS discovery APIs.
POST /api/admin/discover/{region}/auto-register Scan a single region for models, register new ones, update regions for existing
GET /api/admin/discover/{region} Discover available Bedrock models in a region (raw listing)
GET /api/admin/templates Get all 28 editable prompt templates
PATCH /api/admin/templates/{name} Update a template (validates required variables)
POST /api/admin/templates/{name}/reset Reset a template to default
POST /api/admin/templates/{name}/enhance Enhance a template with AI
System
POST /api/log Client-side error/warning logging (recorded as [CLIENT] in server console)
GET /api/health Health check + AWS credential/Bedrock validation

📌 10. Project Structure

ArtSmoker/
├── backend/
│   ├── main.py              # FastAPI app, startup validation, static mount
│   ├── config.py            # Settings (AWS regions, model IDs, paths, limits)
│   ├── model_registry.json  # Single source of truth: models, regions, pricing, format families, quality tiers
│   ├── requirements.txt
│   ├── prompt_templates.json # Editable LLM directive prompts — runtime source of truth (28 templates)
│   ├── routers/
│   │   ├── generate.py      # Two-level asset generation + SSE streaming + reference-guided
│   │   ├── generate_3d.py   # Image-to-3D: generation, source review, variants, async jobs
│   │   ├── custom_deploy.py # Self-hosted model deploy/teardown (Amazon SageMaker)
│   │   ├── styles.py        # Style profile CRUD + directory/S3 import + analysis
│   │   ├── gallery.py       # Asset browsing + file serving + engine exports + bulk delete
│   │   ├── typestudio.py    # Type Studio: text overlay, font serving, AI layout
│   │   ├── video.py         # Video generation (async), job polling, MP4/thumbnail serving
│   │   ├── chat.py          # Chat Studio: LLM streaming, sessions, export, context compaction
│   │   ├── browse.py        # Server-side file/S3 browser for reference import
│   │   ├── refine.py        # Prompt refinement preview + translation preview
│   │   ├── transcribe.py    # Voice transcription
│   │   └── admin.py         # Model registry management + Bedrock discovery + prompt templates
│   ├── services/
│   │   ├── bedrock_client.py     # Shared Bedrock client with connection pooling
│   │   ├── model_registry.py     # Model registry: loads/saves model_registry.json
│   │   ├── prompt_engineer.py    # Claude: prompt refinement + concept generation
│   │   ├── image_generator.py    # Routes to Bedrock (SD 3.5 / Ultra / Core) or SageMaker (FLUX / Hunyuan / Qwen)
│   │   ├── style_analyzer.py     # Two-phase style analysis (cohesion + full)
│   │   ├── post_processor.py     # Stability AI: bg removal, upscale; vtracer: SVG
│   │   ├── transcriber.py        # Nova Sonic: streaming speech-to-text
│   │   ├── import_dedup.py       # Smart deduplication (rotations, animations, folders)
│   │   ├── texture_extractor.py  # glTF/GLB texture extraction
│   │   ├── prompt_translator.py  # Auto-detect language + translate to English
│   │   ├── prompt_templates.py   # Editable LLM directive prompts (load/save/validate)
│   │   ├── video_generator.py   # Video: async Bedrock invoke, S3 download, ffmpeg thumbnails
│   │   ├── cost_tracker.py      # Request-scoped cost accumulator
│   │   ├── custom_models.py    # Self-hosted model catalog (extensible)
│   │   ├── async_jobs.py       # Async generation job queue (Pending Jobs panel)
│   │   ├── reference_analyzer.py # "Inspired by the reference" vision analysis
│   │   ├── mesh_export.py      # Engine-ready 3D exports via managed headless Blender
│   │   ├── sagemaker_deployer.py # Amazon SageMaker endpoint management (direct HF pull for HF models)
│   │   └── sagemaker_invoker.py  # Routes inference to Amazon SageMaker endpoints
│   ├── models/
│   │   ├── style_profile.py       # StyleProfile, AnalyzedStyle, Create/Update
│   │   ├── generation_request.py  # GenerationRequest, AssetType, ImageModel enums
│   │   └── generation_result.py   # GenerationResult, OptionResult, VariantResult
│   └── storage/
│       └── local_store.py         # Local filesystem (S3-compatible interface)
├── frontend/
│   ├── index.html           # SPA entry point
│   ├── css/styles.css       # Dark theme + animations
│   └── js/
│       ├── app.js               # SPA router + DOM caching + navigation + showConfirm()
│       ├── i18n/
│       │   ├── i18n.js          # Core: t() function, language switching, reverse lookup
│       │   ├── en.json          # English (base) — ~1,500 keys
│       │   └── ja/zh/ko/hi/ru/fr/es/de.json   # 8 translations
│       ├── services/api.js      # Backend API client
│       └── components/
│           ├── ImageStudio.js   # 2D Image Studio (options × variations, multi-model)
│           ├── PromptEditor.js  # Two-area prompt editor + compose
│           ├── PromptDesigner.js # Structured prompt decomposition (lock/vary fields)
│           ├── ReferenceStudio.js # Image Inspiration tab (reference-guided generation)
│           ├── TypeStudio.js    # Type Studio (text overlay)
│           ├── VideoStudio.js   # Video Studio (text-to-video generation)
│           ├── ChatStudio.js    # Chat Studio (multi-model LLM chat)
│           ├── Gallery.js       # Gallery grid + search + bulk ops
│           ├── StyleLibrary.js  # Style management + file browser
│           ├── AssetViewer.js   # Full-size preview, edit, 3D, exports, metadata
│           ├── ModelSettings.js # Model registry admin UI (modal)
│           └── VoiceInput.js    # MediaRecorder + transcription
├── data/
│   ├── styles/              # Style profiles + reference images (symlinked)
│   ├── generated/           # Output assets (PNG + SVG + metadata.json)
│   ├── video/               # Video assets (MP4 + thumbnails + job metadata)
│   └── chat/                # Chat sessions (JSON per session)
├── tools/                   # Managed portable Blender (auto-downloaded; gitignored)
├── SPEC.md                  # Full technical specification (rebuild blueprint)
└── README.md                # This file

📌 11. Configurable Limits

Settings in backend/config.py can be overridden via environment variables (prefix ARTSMOKER_):

Setting Env Variable Default Purpose
max_reference_images ARTSMOKER_MAX_REFERENCE_IMAGES 100 Max images imported per style
max_analysis_images ARTSMOKER_MAX_ANALYSIS_IMAGES 20 Max images sent to AI per analysis call
aws_region_models ARTSMOKER_AWS_REGION_MODELS us-west-2 Region for Claude + Stability AI models
aws_region_images ARTSMOKER_AWS_REGION_IMAGES us-east-1 Region for Amazon (Nova Sonic voice, Nova Reel video)
aws_profile ARTSMOKER_AWS_PROFILE None AWS profile name (uses default chain if unset)
auto_update ARTSMOKER_AUTO_UPDATE true Version-gated update on startup + 24h periodic check (git or tarball), then in-place restart

Reducing max_analysis_images reduces AI vision costs per analysis. Reducing max_reference_images limits storage. Both can be tuned based on budget.

📌 12. Amazon Bedrock Pricing & Cost Breakdown

Important

Models deprecate and change fast. New models launch and old ones are retired frequently, so any specific model name or price hard-coded in docs goes stale quickly. ArtSmoker handles this automatically — each Sync from AWS re-discovers the current model lineup, auto-rolls the shared LLM slots to the newest Claude Sonnet/Opus, and refreshes live per-model pricing from the AWS Pricing API into model_registry.json. The app is the source of truth for both which models exist and what they cost (shown live in the Image Studio sidebar per your selected model, quality tier, region, and batch size). Model names and any numbers below are illustrative examples only — always confirm current models/pricing in the app or on the official Amazon Bedrock Pricing page.

The app's default regions are us-west-2 (Claude, Stability AI) and us-east-1 (Amazon Nova Sonic, Nova Reel); prices differ by region. See also SPEC.md for the cost model.

📝 12.1 Per-Unit Pricing

What incurs cost, and its billing unit (see the app for the current per-unit price):

Service Billed Notes
LLM prompt engineering & chat (Claude Sonnet / Opus, auto-rolled to newest on Sync) per input / output token Prompt refinement, concepts, chat, style analysis, moderation
Bedrock image generation (Stable Diffusion 3.5 Large, Stable Image Ultra, Stable Image Core) per image Ultra ≫ SD 3.5 ≫ Core in price; live figure shown in-app
Self-hosted image / 3D (FLUX, HunyuanImage, Qwen-Image, TripoSG, TRELLIS.2) per GPU-second of your SageMaker instance Scale-to-zero when idle ($0); not billed per image
Post-processing (Remove Background, Creative Upscale) per image Stability AI services
SVG conversion free Local (vtracer/potrace) — $0.00

Note

Prices from the official Amazon Bedrock Pricing page as of March 2026. Prices may change — always verify against the official source before budgeting.

📝 12.2 Additional LLM Costs (Per Use)

These LLM calls are included in the generation workflow but not separately itemized in the batch cost tables below:

Call Model When Approx. Cost
Prompt Pre-Check Claude Sonnet Before generation (if toggle enabled) ~$0.005
Moderation Rewrite Claude Sonnet Only when all models reject a prompt ~$0.005
Type Studio Layout Claude Opus Each AI layout suggestion request ~$0.02–$0.05

These are small — pre-check and moderation rewrite are a fraction of a cent each. Type Studio layout is comparable to a single-option prompt refinement.

📝 12.3 Style Analysis Cost (One-Time per Style)

~$0.14 per style (20 images sent to Claude Opus + 8 images cohesion check at Claude Sonnet). The cohesion check adds ~$0.01 (Sonnet with 8 images is very cheap).

📝 12.4 Generation Cost by Batch Size

Includes prompt refinement/concept generation + image generation:

Scenario Stable Image Core Stable Diffusion 3.5 Large Stable Image Ultra
1 option × 1 variation ~$0.05 ~$0.09 ~$0.15
1 option × 5 variations ~$0.21 ~$0.41 ~$0.71
5 options × 5 variations ~$1.05 ~$2.05 ~$3.55

Self-hosted SageMaker models (FLUX, HunyuanImage, Qwen-Image) bill by GPU time on your own instance (scale-to-zero when idle), not per image — see SPEC.md for the compute-cost model.

📝 12.5 Post-Processing Add-Ons (Per Image)

Add-on Per image 1 image 5 images 25 images
Remove Background $0.07 $0.07 $0.35 $1.75
Creative Upscale $0.60 $0.60 $3.00 $15.00
Convert to SVG $0.00 $0.00 $0.00 $0.00

Tip

Creative Upscale note: Handles Stability AI's 16MB response payload limit automatically by using JPEG output format internally, then converting back to PNG. Includes retry with exponential backoff for API throttling.

📝 12.6 Worked Examples

Example Configuration Total Cost
Cheapest 1×1, Stable Image Core, no processing ~$0.05
Standard 1×5, Stable Diffusion 3.5 Large, Remove BG ~$0.76
Full exploration 5×5, Stable Diffusion 3.5 Large, Remove BG + SVG ~$3.80
Premium 5×5, Stable Image Ultra, Remove BG + Upscale + SVG ~$20.30

Tip

Key takeaway: Image generation itself is cheap ($0.01–$0.14/image). Creative Upscale at $0.60/image is the dominant cost — use it selectively on your final chosen assets, not the full batch. Remove Background at $0.07/image is reasonable. SVG conversion is free (runs locally).

📌 13. Disclaimer

Important

Generated Content Quality: All images, videos, and other assets generated by ArtSmoker are produced by AI models available through Amazon Bedrock, including both first-party AWS models and third-party models. The quality, accuracy, and appropriateness of generated content depend entirely on the prompts provided, the models selected, and the style references uploaded by the user. The authors and contributors of ArtSmoker make no guarantees regarding the quality, suitability, or fitness for purpose of any generated content.

Intellectual Property: Users are solely responsible for ensuring that their prompts, reference images, and generated outputs do not infringe on any third-party intellectual property rights, including but not limited to copyrights, trademarks, and personality rights. ArtSmoker is a tool — it does not filter, validate, or assess the IP status of inputs or outputs. The tool authors and contributors bear no responsibility for any IP infringement arising from the use of this software.

AI Model and Service Terms: Generated content is subject to the terms of service and acceptable use policies of the underlying AI model providers accessible through Amazon Bedrock. Users should review the AWS Service Terms, the Amazon Bedrock SLA, and the individual model provider terms before using generated assets in production or commercial contexts.

Model Licensing & Commercial Use: Self-hosted models deployed through ArtSmoker are governed by their creators' license terms, which bind you directly. ArtSmoker surfaces each model's license and dependency breakdown at deploy time and records your acceptance, but it does not verify, enforce, or guarantee your commercial entitlement — staying within license terms (revenue/user thresholds, territory restrictions, attribution requirements, usage reporting) is solely your responsibility. The commercial-licensing guidance in section 1.9.3 is informational only, reflects vendor terms at the time of writing, and is not legal advice; licensing terms change frequently — always confirm the vendor's current terms, and consult counsel before commercial launch. ArtSmoker has no affiliation with, and receives nothing from, any model vendor.

Cost Estimates Only — Monitor Your Own Spend: All costs shown in ArtSmoker (per-image, per-video, per-token, 3D compute, deployment, and session/asset totals) are estimates for guidance only, computed from AWS published pricing and expected usage. They are not a bill and not a guarantee of your actual charges. Real costs depend on your AWS account pricing, region, discounts, taxes, data transfer, endpoint uptime (including idle/warm SageMaker instances), auto-scaling behaviour, and factors outside this tool. You are solely responsible for monitoring and controlling your own AWS spend — use the AWS Billing Console, AWS Cost Explorer, and Budgets/billing alarms to track and cap actual charges. Self-hosted SageMaker endpoints in particular continue to bill for as long as they are deployed or kept warm, even when idle — remember to tear them down when finished. The authors and contributors bear no responsibility for any AWS charges incurred through the use of this software.

No Warranty: This software is provided "as is" without warranty of any kind. See LICENSE for full terms.

📌 14. Full Specification

See SPEC.md for the complete technical specification — architecture, component design, model configuration, API reference, security model, pricing, deployment roadmap, and enough detail to rebuild the project from scratch.

About

A simple, powerful tool for creative teams in game & media studios - to turn concepts & ideas into game-engine-ready 2D art, fully-textured 3D models, and video, with AI doing the heavy lifting while you keep full creative control.

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages