Skill

Generate and Edit Images with AI

Auto-routes image requests between Gemini for realistic people photos and Stable Diffusion for art and edits.


91
Spark score
out of 100
Updated 5 days ago
Version 15.8.0
Models
gemini 2 0

Add to Favorites

Why it matters

Automate the creation, editing, and enhancement of images by intelligently routing requests between specialized AI models for realistic human photos and artistic illustrations.

Outcomes

What it gets done

01

Generate hyper-realistic human photos using ai-studio-image.

02

Create art, illustrations, and edits with Stability AI.

03

Automatically upscale images and remove backgrounds.

04

Perform inpainting and object replacement within images.

Install

Add it to your toolbox

Run in your project directory:

curl -fsSL https://spark.entire.vc/get/ag-image-studio | bash

Overview

IMAGE-STUDIO: Intelligent Image Generator

Image-studio detects what kind of image a request needs and routes it to the right model - Gemini 2.0 Flash for humanized photos of people, Stable Diffusion 3.5 for art, illustration, editing, and upscaling - building an optimized prompt for whichever one it picks. Use it for realistic people photos, illustration, image editing, or upscaling - not for tasks outside image generation or where a more specific tool already fits.

What it does

Image-studio acts as a creative-director router that automatically detects what kind of image a request calls for and picks between two underlying models: ai-studio-image (Gemini 2.0 Flash) for humanized, realistic photos of people, and stability-ai (Stable Diffusion 3.5 Large) for art, illustration, and image editing - so generation, editing, upscaling, background removal, inpainting, and realistic people photos all run through one workflow. A decision matrix drives the routing: a realistic photo of a person or influencer goes to ai-studio-image; illustration, art, or drawing requests go to stability-ai's generate, ultra, or core modes; edits to an existing image go to stability-ai's img2img, inpaint, search-replace, or erase modes; and upscaling or background removal goes to stability-ai's upscale or remove-bg modes - anything ambiguous prompts a clarifying question instead of guessing. ai-studio-image is free (gemini-2.0-flash-exp), applies five layers of narrative humanization (device, lighting, imperfection, authenticity, environment), and ships 20 pre-configured templates (10 influencer-style, 10 educational), but is capped at one image roughly every 9 seconds, about 1K resolution, no custom aspect ratio, and 50 images per day on the free tier. Stability-ai supports 15 styles (photorealistic, anime, digital-art, oil-painting, watercolor, pixel-art, 3d-render, concept-art, comic, minimalist, fantasy, sci-fi, sketch, pop-art, noir) under a Community License that consumes credits, and is not specialized for realistic photos of people. For every request the skill analyzes the image type and goal, selects the ideal model from the decision matrix, builds a prompt optimized for that model, generates with the correct parameters, and presents the result with metadata plus offered variations. If ai-studio-image fails (daily limit or API error), it retries on stability-ai's ultra mode with an adapted prompt and tells the user the model changed; if stability-ai fails (insufficient credits), it retries on ai-studio-image, or if the same image type isn't supported there either, it guides the user on recharging credits. If both models fail, it hands back a detailed prompt the user can run manually and suggests DALL-E, Midjourney, or Leonardo AI as alternatives.

When to use - and when NOT to

Use it for realistic influencer or lifestyle photos of people, professional headshots, humanized educational imagery, illustration or concept art, editing an existing image (inpainting, object replace or erase, background removal), or upscaling an image. Do not use it for tasks unrelated to image generation, when a simpler and more specific tool already covers the request, or when the user just needs general-purpose assistance with no domain-specific image expertise required.

Inputs and outputs

Input is a natural-language image request - subject, action, environment, lighting, and human detail for photo-style prompts; subject, action, artistic style, cinematic lighting, quality descriptors, reference artist, and color palette for stability-ai prompts. Output is the generated image plus its metadata (model used, mode, generation time, save path, dimensions, file size), the optimized prompt actually used, and offered variations such as an alternate-model version or a style or lighting adjustment.

python generate.py --template "instagram-lifestyle" --customization "cafe, manha, sorriso"

Integrations

Routes between the ai-studio-image skill (Gemini 2.0 Flash) and the stability-ai skill (Stable Diffusion 3.5 Large, Community License), and is designed to complement the comfyui-gateway skill; when both primary models fail it falls back to suggesting external tools - DALL-E, Midjourney, or Leonardo AI.

Who it's for

Users who need a single entry point for both realistic people photography and illustrative or artistic image generation - social content creators, marketers producing Instagram or LinkedIn-style visuals, and anyone who wants automatic model selection instead of manually choosing and prompting separate image tools.

FAQ

Common questions

Discussion

Questions & comments · 0

Sign In Sign in to leave a comment.