Explore 240 AI Models
Browse, compare, and integrate the best AI models for video, image, music, audio, and text generation — all through one unified API.
All models available · Real-time pricing
# Available Endpoints
POST /v1/chat/completions # LLM
POST /v1/images/generations # Image
POST /api/v1/kling/text_to_video # Video
POST /v1/audio/speech # TTS
POST /v1/audio/transcriptions # STT
POST /v1/suno/text_to_music # MusicLeading AI models, one API
Claude
Anthropic
Claude API access for Anthropic's LLM across complex reasoning, code, analysis, and extended-context tasks.
ElevenLabs
ElevenLabs
ElevenLabs API access for voice synthesis, text-to-speech, sound effects, speech-to-text, and audio isolation.
Flux
Black Forest Labs
Flux image generation with Dev, Pro, and 2 Klein, plus single-image editing with Dev and Pro.
GPT Image 2
OpenAI
GPT Image 2 adds sharper text rendering, layout control, and consistent multi-image generation.
Kling
Kuaishou
Kling video generation for cinematic text-to-video, image-to-video, motion control, and avatar clips.
Nano Banana
Nano Banana is Google's fast image editor for natural-language edits that keep subjects intact.
Seedance
Bytedance
Seedance 2.5 delivers precise choreography and camera control for text- and image-driven video.
Suno
Suno
Suno v5.5 generates full songs with vocals, lyrics, and stems — no official API available elsewhere.
Veo 3.1
Google Veo 3.1 for native audio-video generation, extension, and upscaling with strong prompt adherence.
DeepSeek
DeepSeek
DeepSeek API access via CAPI — flash for fast, low-cost work; pro for complex agentic tasks.
Embedding
OpenAI
OpenAI text embeddings for semantic search, retrieval, clustering, and ranking workflows.
Fish Audio
Fish Audio
Fish Audio API access for expressive multilingual and production-grade text-to-speech with managed MP3 or WAV output.
Flux 2
Black Forest Labs
Flux 2 API access for text-to-image and remix-image with strong prompt adherence from Black Forest Labs.
Flux Kontext
Black Forest Labs
Flux Kontext API access for in-context image editing, local edits, style transfer, and character consistency.
Gemini
Gemini API access for Google's multimodal LLM across chat, code generation, reasoning, and long-context tasks.
Gemini Omni
Gemini Omni API access for voice, character, and multimodal video resources in agent media workflows.
Gemini TTS
Gemini TTS API access for multi-speaker dialogue with configurable voices, accents, delivery styles, and pacing.
GLM
Z.ai
Z.ai GLM API access via CAPI — MIT-licensed MoE models with up to 200K context, leading open-weight coding benchmarks.
GPT
OpenAI
OpenAI's flagship reasoning and chat models through the OpenAI-compatible chat completions endpoint.
GPT Image
OpenAI
OpenAI GPT Image for instruction-following generation and conversational image editing.
Grok
xAI
xAI Grok models for real-time reasoning, coding, and tool-driven agent workflows.
Grok Imagine
xAI
Grok Imagine turns short prompts into stylised stills with quick turnaround and playful aesthetic control.
Hailuo
MiniMax
MiniMax Hailuo video models for expressive character motion and physics-aware rendering.
HappyHorse
Alibaba
HappyHorse video models optimised for fast draft iterations and edit-in-place workflows.
Ideogram V3
Ideogram
Ideogram V3 is built for legible in-image typography, posters, logos, and design mockups.
Imagen 4
Google Imagen 4 for photorealistic rendering with accurate lighting and material detail.
InfiniteTalk
MeiGen-AI
InfiniteTalk turns a single portrait plus audio into a talking-head video with lip sync.
Kimi
Moonshot AI
Moonshot AI Kimi models tuned for long-document reading, research synthesis, and agentic search.
Lip Sync
Bytedance
Volcengine lip sync re-dubs an existing video to new audio while preserving the original performance.
Luma
Luma
Luma Ray models for smooth motion synthesis, keyframe control, and video modification.
Midjourney
Midjourney
Midjourney v7 through the API — stylised generation with reference and character consistency.
MiMo
Xiaomi
Xiaomi MiMo lightweight reasoning models for high-throughput classification and extraction.
MiniMax H3
MiniMax
MiniMax H3 for high-fidelity long-form generation with multi-shot consistency.
OmniHuman
Bytedance
OmniHuman audio-to-video avatars with subject detection and human identification helpers.
OpenAI TTS
OpenAI
OpenAI text-to-speech with low-latency streaming voices for real-time assistants.
PixVerse
PixVerse
PixVerse for stylised video generation, transitions, extension, and clip editing.
Producer
Producer
Producer creates loopable beds, stems, and adaptive music cues for games and video.
Qwen
Alibaba
Alibaba Qwen text models offering strong multilingual reasoning at open-weight pricing.
Qwen Image
Alibaba
Qwen Image supports multilingual prompts with precise text rendering and layout control.
Recraft
Recraft
Recraft generates vector art, icon sets, and brand-consistent illustrations for product teams.
Runway
Runway
Runway Gen-4 and Aleph for text-to-video generation and instruction-based video editing.
Seedream
Bytedance
Seedream produces high-resolution, prompt-faithful images with strong typography.
Topaz Upscale
Topaz
Topaz upscaling restores detail and clean edges in existing video at up to 4× resolution.
Transcription
OpenAI
Whisper-based speech-to-text with timestamps, diarisation, and translation for long recordings.
Video Utility
Topaz
Frame interpolation, matting, and background removal helpers for post-production pipelines.
Wan Image
Alibaba
Alibaba Wan image models for foundational text-to-image and Chinese-typography rendering.
Wan Video
Alibaba
Alibaba Wan video models for text-to-video, image-to-video, speech-driven avatars, and editing.
Z-Image
Z.ai
Z-Image is a compact, fast generator tuned for low-latency previews and bulk batches.
Not sure which model to pick?
Every model page lists the exact model IDs, per-unit pricing, and a runnable request you can copy.