Model Catalog

Explore 240 AI Models

Browse, compare, and integrate the best AI models for video, image, music, audio, and text generation — all through one unified API.

All models available · Real-time pricing

REST API
# Available Endpoints

POST   /v1/chat/completions         # LLM
POST   /v1/images/generations       # Image
POST   /api/v1/kling/text_to_video  # Video
POST   /v1/audio/speech             # TTS
POST   /v1/audio/transcriptions     # STT
POST   /v1/suno/text_to_music       # Music

Leading AI models, one API

ElevenLabsKlingGPT Image 2Veo 3.1HailuoLumaPixVerseRecraftFish Audio
48 models
Provider
Text

Claude

Anthropic

Claude API access for Anthropic's LLM across complex reasoning, code, analysis, and extended-context tasks.

from $0.0006 / 1K tokensView
Audio

ElevenLabs

ElevenLabs

ElevenLabs API access for voice synthesis, text-to-speech, sound effects, speech-to-text, and audio isolation.

from $0.040 / 1K charsView
Image

Flux

Black Forest Labs

Flux image generation with Dev, Pro, and 2 Klein, plus single-image editing with Dev and Pro.

from $0.030 / callView
Image

GPT Image 2

OpenAI

GPT Image 2 adds sharper text rendering, layout control, and consistent multi-image generation.

from $0.030 / callView
Video

Kling

Kuaishou

Kling video generation for cinematic text-to-video, image-to-video, motion control, and avatar clips.

from $0.070 / secondView
Image

Nano Banana

Google

Nano Banana is Google's fast image editor for natural-language edits that keep subjects intact.

from $0.025 / callView
Video

Seedance

Bytedance

Seedance 2.5 delivers precise choreography and camera control for text- and image-driven video.

from $0.090 / secondView
Music

Suno

Suno

Suno v5.5 generates full songs with vocals, lyrics, and stems — no official API available elsewhere.

from $0.180 / callView
Video

Veo 3.1

Google

Google Veo 3.1 for native audio-video generation, extension, and upscaling with strong prompt adherence.

from $0.150 / secondView
Text

DeepSeek

DeepSeek

DeepSeek API access via CAPI — flash for fast, low-cost work; pro for complex agentic tasks.

from $0.0004 / 1K tokensView
Text

Embedding

OpenAI

OpenAI text embeddings for semantic search, retrieval, clustering, and ranking workflows.

from $0.0000 / 1K tokensView
Audio

Fish Audio

Fish Audio

Fish Audio API access for expressive multilingual and production-grade text-to-speech with managed MP3 or WAV output.

from $0.0000 / 1K UTF-8 bytesView
Image

Flux 2

Black Forest Labs

Flux 2 API access for text-to-image and remix-image with strong prompt adherence from Black Forest Labs.

from $0.070 / callView
Image

Flux Kontext

Black Forest Labs

Flux Kontext API access for in-context image editing, local edits, style transfer, and character consistency.

from $0.110 / callView
Text

Gemini

Google

Gemini API access for Google's multimodal LLM across chat, code generation, reasoning, and long-context tasks.

from $0.0003 / 1K tokensView
Video

Gemini Omni

Google

Gemini Omni API access for voice, character, and multimodal video resources in agent media workflows.

from $0.0000 / callView
Audio

Gemini TTS

Google

Gemini TTS API access for multi-speaker dialogue with configurable voices, accents, delivery styles, and pacing.

from $0.0014 / 1K tokensView
Text

GLM

Z.ai

Z.ai GLM API access via CAPI — MIT-licensed MoE models with up to 200K context, leading open-weight coding benchmarks.

from $0.0001 / 1K tokensView
Text

GPT

OpenAI

OpenAI's flagship reasoning and chat models through the OpenAI-compatible chat completions endpoint.

from $0.0005 / 1K tokensView
Image

GPT Image

OpenAI

OpenAI GPT Image for instruction-following generation and conversational image editing.

from $0.020 / callView
Text

Grok

xAI

xAI Grok models for real-time reasoning, coding, and tool-driven agent workflows.

from $0.0003 / 1K tokensView
Image

Grok Imagine

xAI

Grok Imagine turns short prompts into stylised stills with quick turnaround and playful aesthetic control.

from $0.020 / callView
Video

Hailuo

MiniMax

MiniMax Hailuo video models for expressive character motion and physics-aware rendering.

from $0.060 / secondView
Video

HappyHorse

Alibaba

HappyHorse video models optimised for fast draft iterations and edit-in-place workflows.

from $0.030 / secondView
Image

Ideogram V3

Ideogram

Ideogram V3 is built for legible in-image typography, posters, logos, and design mockups.

from $0.040 / callView
Image

Imagen 4

Google

Google Imagen 4 for photorealistic rendering with accurate lighting and material detail.

from $0.040 / callView
Video

InfiniteTalk

MeiGen-AI

InfiniteTalk turns a single portrait plus audio into a talking-head video with lip sync.

from $0.180 / callView
Text

Kimi

Moonshot AI

Moonshot AI Kimi models tuned for long-document reading, research synthesis, and agentic search.

from $0.0002 / 1K tokensView
Video

Lip Sync

Bytedance

Volcengine lip sync re-dubs an existing video to new audio while preserving the original performance.

from $0.120 / callView
Video

Luma

Luma

Luma Ray models for smooth motion synthesis, keyframe control, and video modification.

from $0.055 / secondView
Image

Midjourney

Midjourney

Midjourney v7 through the API — stylised generation with reference and character consistency.

from $0.050 / callView
Text

MiMo

Xiaomi

Xiaomi MiMo lightweight reasoning models for high-throughput classification and extraction.

from $0.0001 / 1K tokensView
Video

MiniMax H3

MiniMax

MiniMax H3 for high-fidelity long-form generation with multi-shot consistency.

from $0.080 / secondView
Video

OmniHuman

Bytedance

OmniHuman audio-to-video avatars with subject detection and human identification helpers.

from $0.200 / callView
Audio

OpenAI TTS

OpenAI

OpenAI text-to-speech with low-latency streaming voices for real-time assistants.

from $0.015 / 1K charsView
Video

PixVerse

PixVerse

PixVerse for stylised video generation, transitions, extension, and clip editing.

from $0.045 / secondView
Music

Producer

Producer

Producer creates loopable beds, stems, and adaptive music cues for games and video.

from $0.090 / callView
Text

Qwen

Alibaba

Alibaba Qwen text models offering strong multilingual reasoning at open-weight pricing.

from $0.0002 / 1K tokensView
Image

Qwen Image

Alibaba

Qwen Image supports multilingual prompts with precise text rendering and layout control.

from $0.020 / callView
Image

Recraft

Recraft

Recraft generates vector art, icon sets, and brand-consistent illustrations for product teams.

from $0.040 / callView
Video

Runway

Runway

Runway Gen-4 and Aleph for text-to-video generation and instruction-based video editing.

from $0.120 / secondView
Image

Seedream

Bytedance

Seedream produces high-resolution, prompt-faithful images with strong typography.

from $0.025 / callView
Utility

Topaz Upscale

Topaz

Topaz upscaling restores detail and clean edges in existing video at up to 4× resolution.

from $0.050 / callView
Audio

Transcription

OpenAI

Whisper-based speech-to-text with timestamps, diarisation, and translation for long recordings.

from $0.006 / minuteView
Utility

Video Utility

Topaz

Frame interpolation, matting, and background removal helpers for post-production pipelines.

from $0.020 / callView
Image

Wan Image

Alibaba

Alibaba Wan image models for foundational text-to-image and Chinese-typography rendering.

from $0.018 / callView
Video

Wan Video

Alibaba

Alibaba Wan video models for text-to-video, image-to-video, speech-driven avatars, and editing.

from $0.035 / secondView
Image

Z-Image

Z.ai

Z-Image is a compact, fast generator tuned for low-latency previews and bulk batches.

from $0.010 / callView

Not sure which model to pick?

Every model page lists the exact model IDs, per-unit pricing, and a runnable request you can copy.

POST/v1/chat/completions
POST/v1/images/generations
POST/v1/kling/text_to_video