Pricing Blog

Image Models

Unlock unlimited creative possibilities. Our AI image models help you create professional-quality visuals effortlessly – no design skills required.

GPT Image 2.5
GPT Image 2.5
OpenAI

OpenAI’s newest image model in two versions — Flare for speed, Sunburst for maximum detail — with 1K/2K/4K output and xhigh/max quality.

Qwen-Image 3.0
Qwen-Image 3.0
Alibaba

Alibaba's dense-text image flagship — 4,500-token prompts, 12 languages, readable ~10px text, one-pass infographics.

MAI-Image 2.5
MAI-Image 2.5
Microsoft

Microsoft's first-party image model — top-tier editing and text rendering; base tier plus Flash and Pro.

Ideogram 4.0
Ideogram 4.0
Ideogram

Open-weight typography king — best-in-class in-image text, native 2K, JSON layout prompting, transparent output.

Reve 2.1
Reve 2.1
Reve

Native 4K (16MP) layout-first image model — plans structure before rendering; debuted at #2 on the Arena Text-to-Image leaderboard (July 2026).

Midjourney V8.2
Midjourney V8.2
Midjourney

July 2026 default: bolder aesthetics, sharper personalization, native 2K HD, 24× faster style exploration.

GPT Image 2
GPT Image 2
OpenAI

GPT Image 2: OpenAI's April 2026 image model with near-perfect text rendering, strong world knowledge and up to 4K output. Available now on SharkFoto.

Nano Banana 2
Nano Banana 2
Google

Nano Banana 2: Google's AI image generator with Pro-level quality at Flash speed. 4K resolution, subject consistency, precision text, world knowledge.

Wan 2.7 Image
Wan 2.7 Image
Alibaba

Wan 2.7 Image: Alibaba's unified AI model for text-to-image and image editing. Thinking mode, 9-reference inputs, 4K Pro output, 12-language text rendering.

Qwen Image 2.0
Qwen Image 2.0
Alibaba

Qwen Image 2.0: Alibaba's unified AI image model. 1k-token instructions, native 2K resolution, professional PPT/poster generation. 7B diffusion decoder with an 8B Qwen3-VL encoder.

Seedream 5.0
Seedream 5.0
ByteDance

Seedream 5.0 Lite: ByteDance's AI image model. Advanced reasoning, photorealistic visuals, precise editing. 2K or 3K output on SharkFoto for professional creative projects.

GPT Image 1.5
GPT Image 1.5
OpenAI

GPT Image 1.5: OpenAI's December 2025 AI image generator with precise editing, up to 4x faster generation, and advanced text rendering. Create and edit images.

Google Nano Banana
Google Nano Banana
Google

Google Nano Banana: Fast AI image generation with Gemini 2.5 Flash Image. Character consistency, multi-image blending, 1024px resolution.

Google Nano Banana Pro
Google Nano Banana Pro
Google

Nano Banana Pro: Studio-quality AI image generation with clear text, 4K resolution, and Gemini 3 reasoning. Create professional visuals.

Seedream 4.5
Seedream 4.5
ByteDance

Seedream 4.5: ByteDance's AI image model with industry-leading text rendering, multi-image editing, and 4K quality output for professionals.

Video Models

Turn your ideas and images into captivating videos with our state-of-the-art AI video models. Perfect for storytelling, marketing, and creative projects – produce professional-quality videos in minutes.

Wan 3.0
Wan 3.0
Alibaba

Alibaba's newest video model - text-to-video and image-to-video at up to 1080p, any length from 5 to 15 seconds, native audio on or off.

Runway Gen-4.5
Runway Gen-4.5
Runway

Runway's frontier video model — cinematic text-to-video at 720p, launch-day AA leaderboard #1.

Luma Ray 3.2
Luma Ray 3.2
Luma

Dream Machine video with frame-level control — up to 16 keyframes, 1080p output, up to 20s in Modify Video, native 16-bit HDR export.

MiniMax H3 (Hailuo 3.0)
MiniMax H3 (Hailuo 3.0)
MiniMax

Omni-modal video — 2K clips up to 15s, native stereo audio, omni-reference from up to 12 inputs.

Happy Horse 1.0
Happy Horse 1.0
Alibaba

Alibaba's #1-ranked video model — 1080p with native synchronized audio and multilingual lip-sync in one pass.

Gemini Omni Flash
Gemini Omni Flash
Google

Conversational video generation and editing — 720p, native audio, edit by talking, $0.10 per second.

FLUX 3
FLUX 3
Black Forest Labs

Multimodal frontier model — FLUX 3 Video makes clips up to 20 seconds long with optional native audio, from text or images.

Seedance 2.5
Seedance 2.5
ByteDance

Native 30-second clips in one continuous take — up to 50 image, video and audio references, video editing, and unified audio-video.

Seedance 2.0
Seedance 2.0
ByteDance

Seedance 2.0: ByteDance's cinematic AI video model. 15-file multi-modal reference, up to 1080p on SharkFoto, native audio, one-sentence editing. Pro-level video creation.

Kling O3
Kling O3
Kuaishou

Kling O3 (Video 3.0 Omni): Kuaishou's next-gen unified multimodal video model. Up to 15s clips with native multilingual audio, multi-shot storyboards and reference-based consistency.

Vidu Q3
Vidu Q3
Vidu

Vidu Q3: Industry's first 16-second native audio-video AI model. Smart Cuts, cinematic camera control, multi-shot storytelling, 1080p output. Ranked #2 globally on Artificial Analysis at launch.

Seedance 1.5 Pro
Seedance 1.5 Pro
ByteDance

Seedance 1.5 Pro: ByteDance's audio-visual generation model with film-grade cinematography, native audio, and powerful storytelling capabilities.

Wan 2.6
Wan 2.6
Alibaba

Wan 2.6: Alibaba's multimodal AI video model with native audio, multi-shot storytelling, and 1080p cinematic quality up to 15 seconds.

Kling O1
Kling O1
Kuaishou

Kling O1: World's first unified multimodal video model. Input anything, understand everything. 3-10s flexible duration with industrial-grade consistency.

Kling VIDEO 2.6 Pro
Kling VIDEO 2.6 Pro
Kuaishou

Kling VIDEO 2.6 Pro: Kling's first native audio video model. Generate complete audio-visual videos with voiceovers, sound effects, and ambient atmosphere.

Runway Gen-4 Aleph
Runway Gen-4 Aleph
Runway

Runway Gen-4 Aleph: Runway's 2025 in-context video editing model, now succeeded by Aleph 2.0. For prompt-based video-to-video editing on SharkFoto, try Seedance 2.5.

Google Veo 3.1
Google Veo 3.1
Google

Google Veo 3.1: State-of-the-art AI video generation with native audio, 720p/1080p output on SharkFoto (model supports up to 4K), and advanced creative controls. Best-in-class performance.

OpenAI Sora 2
OpenAI Sora 2
OpenAI

OpenAI Sora 2: OpenAI's 2025 video and audio generation model, retired from its API on September 24, 2026. For AI video with native audio on SharkFoto, try Google Veo 3.1.

PixVerse V5
PixVerse V5
PixVerse

PixVerse V5: AISphere's AI video model with ultra-resolution engine, cinematic camera control, and fusion features. Top-ranked on Artificial Analysis at launch.

Audio Models

Explore our collection of AI music and audio generators designed for musicians, content creators, and producers. Create original tracks, enhance audio quality, and bring your sonic ideas to life.