Express Computer
Home  »  News  »  ZEE5 and Google Cloud Collaborate to Reimagine Content Discovery with ‘Clips’

ZEE5 and Google Cloud Collaborate to Reimagine Content Discovery with ‘Clips’

Built with Google Cloud, the technology combines Gemini models, multilingual speech AI and computer vision to identify high-engagement moments and automatically transform long-form content into short-form vertical videos at scale

0 0

ZEE5 and Google Cloud are collaborating to transform how long-form entertainment content is processed and surfaced for discovery through Clips, an AI-powered short-form video experience that automatically identifies engaging moments from films and shows and converts them into vertical videos.

The technology brings together Google Cloud’s Gemini models, Chirp multilingual speech model, Cloud Vision API and deep-learning models for image processing to analyse and process large volumes of long-form content. By combining speech, visual and contextual signals, the system identifies moments that are suitable for short-form discovery and automatically packages them into vertical video clips.

This approach enables content discovery assets to be generated from an existing catalogue at scale, significantly reducing the manual effort traditionally required to identify, edit and package individual moments from long-form content. To date, the technology has generated 13,000 Clips from 1,200 titles, demonstrating its ability to process and transform content across a large library.

The technology is designed to create a continuous discovery layer around long-form content. A viewer can encounter a short clip, understand the context and, where it sparks interest, move directly into the corresponding full-length title. This creates a technology-driven pathway between short-form discovery and long-form consumption.

“At ZEE5, we are looking at content consumption differently, and Clips is a direct reflection of that. Developed in collaboration with Google Cloud, it puts compelling moments from our long-form catalogue into the swipe journey. What makes this particularly exciting is the ability to turn thousands of hours of existing content into fresh entry points at scale using AI. The early traction shows us that when technology is applied with a clear understanding of how people consume content today, it can fundamentally change how audiences find their next watch,” said Tejkarran Singh Bajaj, Business Head – ZEE5 India and AI & Innovation.

The underlying technology combines multiple AI capabilities to automate a process that would otherwise require extensive manual intervention. Gemini models help analyse content and identify relevant moments, while Chirp supports multilingual speech processing and Cloud Vision API enables visual analysis. Together with deep-learning models for image processing, these capabilities allow the system to understand and process different elements of long-form video before generating short-form outputs.

“Clips was built around a simple audience need: making it easier for viewers to discover stories they may not have actively searched for. We are using AI to identify engaging moments within our long-form content and automatically transform them into short, vertical videos. The solution brings together Gemini AI, speech-to-text and Cloud Vision API to process and package content at scale. We have generated 13,000 Clips from 1,200 titles so far, demonstrating what this technology can unlock across a large content library,” said Pramod Prakash, Chief Technology Officer, Zee Entertainment Enterprises Ltd.

The collaboration also demonstrates how generative AI can be integrated into production-scale digital workflows to create new product capabilities while working across large and diverse content libraries. Amit Kumar, Managing Director, Digital Natives Business, Google Cloud India, said, “Our collaboration with ZEE5 on Clips is a great example of our deep partnership and demonstrates how generative AI can drive measurable business value at scale. By leveraging Google Cloud’s advanced Gemini models, these joint AI innovations accelerate time-to-market for new platform features while continuously elevating the daily viewer experience.”

By combining multimodal AI, speech intelligence, computer vision and automated content processing, Clips demonstrates how AI can move beyond content generation to analyse existing media libraries, identify contextually relevant moments and create new formats for content discovery at scale.

Leave A Reply

Your email address will not be published.