Persona-1-Live
Live stream and interact with Persona-1.

Today, we're excited to share persona-1-live, a real-time streaming variant of persona-1 — capable of holding dynamic, human-feeling conversations with low latency.
Architecture Overview
persona-1-live enhances persona-1's flow-based image transformer with a next-token prediction scheme optimized for interactive video generation. Crucially, our implementation enables conversational interactivity without sacrificing persona-1's speed, expressive fidelity, or hallmark low cost.

Empirically, we've also found the new next-token prediction scheme exploits our relatively small training sets more effectively, resulting in broader and more nuanced expression diversity. We invite you to video chat with persona-1-live in our playground.
Limitations & Future Work
We've discovered that our current audio encoder performs best with longer windows — much longer than the new next-token prediction scheme likely needs. As a result, the typical delay between the end of a user's turn and the persona's response is roughly six seconds. Experiments with alternative audio encoders are showing promising results, and we expect to bring this delay down by roughly 50% in the near future.
Pricing & Availability
persona-1-live is currently in early access, with pricing comparable to persona-1: just a few cents per minute, ~1/2 the cost of the cheapest competing models.
We're excited to see what sorts of novel applications today's live streaming upgrade to persona-1 enables.