Blog
Research

Persona-1-Live

Live stream and interact with Persona-1.

Persona-1-Live cover art: the model name in gradient brackets over a field of glowing attention cells

Today, we're excited to share persona-1-live, a real-time streaming variant of persona-1 — capable of holding dynamic, human-feeling conversations with low latency.

Architecture Overview

persona-1-live enhances persona-1's flow-based image transformer with a next-token prediction scheme optimized for interactive video generation. Crucially, our implementation enables conversational interactivity without sacrificing persona-1's speed, expressive fidelity, or hallmark low cost.

The new next-token prediction scheme for interactive video generation

Empirically, we've also found the new next-token prediction scheme exploits our relatively small training sets more effectively, resulting in broader and more nuanced expression diversity. We invite you to video chat with persona-1-live in our playground.

Limitations & Future Work

We've discovered that our current audio encoder performs best with longer windows — much longer than the new next-token prediction scheme likely needs. As a result, the typical delay between the end of a user's turn and the persona's response is roughly six seconds. Experiments with alternative audio encoders are showing promising results, and we expect to bring this delay down by roughly 50% in the near future.

Pricing & Availability

persona-1-live is currently in early access, with pricing comparable to persona-1: just a few cents per minute, ~1/2 the cost of the cheapest competing models.

We're excited to see what sorts of novel applications today's live streaming upgrade to persona-1 enables.