What Is Reka? How to Use This Multimodal AI That Understands Images, Video, and Audio [Latest 2026]
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated
@aifriends
AI Friends(https://aifriends.jp)のクロスポスト公式アカウント。AIツールの紹介・使い方・できることを、中学生でもわかるやさしい日本語で届けます。
What you'll learn in this article
Reka is a startup company developing "multimodal AI" that can simultaneously understand text, images, video, and audio. "Multimodal" means the ability to handle multiple types of data at once — Reka's AI can read the content of photos and videos and answer questions about them, not just process written text. In June 2026, the company merged with video-generation AI company Moonvalley and is now advancing the development of AI that understands the physical world (Physical AI) at a higher level. Founded by researchers formerly at DeepMind and Google, the company is attracting attention for its cutting-edge technical capabilities.
Reka offers four models suited to different use cases, ranging from the lightweight Spark (1 billion parameters) to the highest-performing Core (67 billion parameters). The Edge model is notable for being light enough to embed in robots and cameras, while the Flash model operates quickly while delivering performance on par with GPT-3.5 and Gemini Pro. The Core model is the top-tier offering that can compete with GPT-4V and Claude-3 Opus, and excels at deeply understanding video content. It can handle complex tasks such as extracting information from images and video, and answering questions that combine audio and text.
Reka is accessed via an API (a mechanism for use from programs). Simply create an account on the official website, obtain an API key (think of it as a password for connecting), and you can start using it right away. Pricing is pay-as-you-go based on actual usage, with no upfront costs. Charges are based on the amount of input data (number of tokens); the Core model costs approximately $3 per 1 million tokens, while lighter models are available at lower rates. The Vision API, which supports video analysis, comes with a free trial of 3 hours (180 minutes) of usage, so you can try it out before committing to full adoption. Deployment environments are also flexible, with options including cloud, on-premises, and VPC.
Advantages include the fact that text, images, video, and audio are trained together from the ground up, enabling the AI to accurately understand the relationships between different types of data. With four model sizes available, it covers a wide range of use cases — from lightweight applications embedded in smartphones and robots to complex enterprise-level analysis. The new Edge model announced in March 2026 is also a strength, as it can be embedded in physical devices and run in real time. Disadvantages include the limited availability of Japanese-language information, meaning usage is primarily centered around English. Additionally, there is no general-purpose chat service like GPT-4 or Claude for end users, so technical knowledge of working with APIs is required.
Reka is especially recommended for companies and developers who want to automatically analyze the content of images and video. For example, it is well suited for applications such as AI-powered surveillance of security camera footage or robots that assess their surroundings and act accordingly. In fact, Reka's physical security system was showcased at the Smart City Asia event in 2026. It is also a good fit for engineers who want to build applications that handle multiple types of data simultaneously, and for researchers tackling advanced multimodal tasks that existing AI cannot address. Because you can start small with pay-as-you-go pricing, it is also easy for startups to try out when testing new services.
This article is a cross-post from AI Friends.