Multimodal AI that sees, hears, and reads โ in one API. Built by ex-Google DeepMind.
Reka is what happens when the team that helped build Gemini goes independent. Three multimodal models โ Core (frontier), Edge (mid-size), Flash (small) โ that natively understand text, images, audio, and short video clips. If you're building a product that needs vision + language + audio in one model, Reka is the most cohesive multimodal API in 2026.
Who it's for: Developers building video understanding, voice agents with vision, document AI, or any product where one model needs to ingest multiple modalities.
One API handles text, images, audio, and short video clips natively. No need to chain a vision model + speech-to-text + LLM. Reka fuses them in the same forward pass.
Pass a 30-second video clip and ask questions about it. Core can identify actions, objects, on-screen text, and answer follow-ups. Better than Gemini 1.5 Pro on short-form video benchmarks.
128k token context for Core, 64k for Edge. Enough to drop in entire transcripts, long documents, or full meeting recordings.
Core at $0.40/M input tokens โ about 1/3 the price of GPT-5 for comparable tasks. Edge and Flash drop to $0.04/M for high-volume workloads.
Reka is the right pick when your AI workload genuinely spans modalities โ video analysis, voice agents with screen understanding, document AI that includes charts and images. The multimodal fusion is genuinely better than stitching separate APIs. For pure text or pure code, Claude/GPT-5 still win. For vision-only, GPT-5 Vision or Claude Sonnet 4 are simpler. The sweet spot: when one model needs to see, hear, and reason at the same time.
Google's multimodal flagship. 1M+ context window, video, audio. Pricier than Reka.
Best for text reasoning + image understanding. No native audio/video but stronger on code.
GPT-5 with voice, vision, image gen. The generalist. Stronger brand recognition than Reka.
OpenAI's multimodal. Cheaper than Reka Core but no video understanding yet.