TwelveLabs Brings Video Understanding to Physical AI With Latest Launch

TwelveLabs Brings Video Understanding to Physical AI With Latest Launch TwelveLabs Brings Video Understanding to Physical AI With Latest Launch Pegasus' physical AI capabilities turn egocentric video into data for teaching machines to understand the real world GlobeNewswire October 06, 2026

SAN FRANCISCO, Oct. 06, 2026 (GLOBE NEWSWIRE) -- TwelveLabs, a leading video intelligence company, today announced the release of Pegasus 1.6, which adds new capabilities for understanding and navigating complex real-world environments. This latest model enables TwelveLabs expansion into physical AI, the domain of artificial intelligence that can perceive, reason about, and act in the physical world through machines like robots, drones, autonomous vehicles, and industrial equipment. Now TwelveLabs can solve key workflow-specific challenges and unlock the full potential for physical AI.

As the physical world rapidly digitalizes, physical AI teams are gathering vast amounts of complex video, but transforming raw footage into actionable model intelligence remains a critical challenge. TwelveLabs bridges this gap, automatically generating rich insights from real-world perspectives and enabling machines to see, understand, and safely interact with their environments like never before.

"Our mission has always been to help machines understand how the world works through video," said Jae Lee, CEO and co-founder of TwelveLabs. "Physical AI is the next expression of that mission. Most of what people know about doing physical work, such as a changing grip, or a recovery after something slips, has never been captured in a form a machine can learn from. With our newest model release, we can now turn that footage into structured, reviewable knowledge, so robotics and physical AI teams can train on real human experience instead of starting from scratch."

Notably, TwelveLabs’ new Pegasus 1.6 model is the first TwelveLabs model built to understand egocentric video. Egocentric video is shot from the point of view of the person doing the work, whether that’s someone cooking a meal, assembling parts on a factory line, or operating a robot remotely.

Pegasus 1.6 currently supports five core workflows powered by this video-native model. They are:

Pegasus 1.6 builds on the industry-leading video understanding capabilities TwelveLabs developed for enterprises that maintain massive video libraries.

The new release extends Pegasus 1.5’s foundational functionality that attracted several new customers, particularly from Time-Based Metadata (TBM), which allows users to define a custom schema and automatically receive timestamped, structured metadata from video content. This is particularly useful in aiding contextual understanding, as people often narrate what they are doing in egocentric clips. Pegasus 1.6 also improved entity recognition for more consistent tracking of hands, objects, and tools across clips. The model also offers faster, more cost-efficient processing for high-volume video workloads.

To learn more about Pegasus 1.6 and all that it can do for physical AI, please visit https://www.twelvelabs.io/solutions/physical-ai. To find out more about Pegasus 1.6’s performance benchmarks, download its latest whitepaper here.

About TwelveLabs
TwelveLabs makes video one of an organization’s most useful assets. The company has created a full-stack video intelligence platform, powered by its industry leading video understanding models Marengo and Pegasus, designed to see and understand video the way humans do in a fraction of the time. Developers, enterprises, and creatives can now leverage this single system that gets smarter over time to access and act on all of their video content. TwelveLabs is backed by leading investors, including Amazon, Index Ventures, Naver, NEA, and Radical Ventures. For more information, visit www.twelvelabs.io.


Media Contact:
Amber Moore
amber@moorecom2.com

Primary Logo