From source document to finished podcast, in one place.
PRODUCT OVERVIEW
What uStudio Labs is
uStudio Labs builds products, prototypes, and solutions for uStudio customers at speed. Operating outside the core uStudio product organization, the team draws on the uStudio Platform and its podcasting capabilities to turn customer requests into working solutions quickly. Because these are prototypes and beta releases, customer feedback shapes them directly - it is what lets the team iterate and deliver business value sooner. That same pace means Labs products change more often than the core uStudio products and may introduce new features, or the occasional rough edge, without prior notice. To request access or share feedback, email labs@ustudio.com.
What Autocast is
Autocast turns your source material - emails, whitepapers, blog posts, web articles, presentations, and PDFs - into finished, on-brand podcasts in minutes rather than the hours or days traditional production takes. Just as important, it collapses the entire production pipeline into a single workspace.
Instead of coordinating writers, voice talent, recording studios, music libraries, audio engineers, and post-production vendors, one Autocast workspace covers the whole journey: ingest the source, draft the script, cast and voice the hosts, generate and place music, record live guests, mix and master, and publish the episode - with no external studios and no vendor hand-offs.
The experience blends three ideas: the source-to-script simplicity of dropping in documents and getting an episode script, the script-first editing of a screenplay you can direct line by line, and the plain-language workflow of starting to type, dragging in references, and iterating in ordinary English.
Why teams use it
- Speed. Go from raw material to a mixed, mastered episode in minutes, not days.
- One workspace. The whole pipeline lives in a single place - no stitching together music intros and outros in separate tools.
- On-brand and on-message. Episodes are grounded in your own documents, so the conversation reflects your actual content rather than generic filler.
- Editorial control. Direct performances line by line and refine anything in plain English - you stay in charge of the final result.
- Studio-quality output. Automatic mixing, mastering, and a self-checking quality pass deliver clean audio ready to publish.
What you can do with Autocast
Every step below happens inside the same workspace.
| Capability | What it does for you |
|---|---|
| Bring in your material | Upload PDFs, documents, text, and images, or paste in web and article URLs. You can also start from a plain-language prompt with no files at all. Your sources stay bundled with the episode so you can reopen and regenerate against them anytime. |
| Draft the show automatically | Autocast reads and understands your material, then drafts a complete, multi-host episode grounded in that content - first an outline (title, summary, hosts, and chapter structure), then full back-and-forth dialogue. Any chapter can be regenerated on its own. |
| Edit like a script | Dialogue lands in a clean, screenplay-style editor. Click any line to rewrite it, add or remove lines, reassign a line to a different host, and preview any line as speech on the spot. |
| Direct the performance | Add per-line cues - a laugh, a pause, emphasis, a shift in tone - and the AI voices perform them. Set the overall scene and director's notes to shape pacing across the whole episode. |
| Edit in plain English | Describe the change you want (“make the intro punchier,” “add a laugh here,” “add intro music”) and Autocast applies it - across a selected range or the whole episode. Every change is tracked and reversible. |
| Cast your hosts | Create and shape hosts with a name, role, personality, and backstory, and give each a distinct voice from 30 AI voices spanning a range of tones and pitches. Hosts can be AI-voiced or performed by real people. |
| Interview real people, live | An AI host can hold a real-time, natural conversation with a real guest in the recording booth - taking turns automatically - and that take is folded straight into the episode. |
| Record real voices | Capture a full multi-participant session in one take, invite remote guests through a one-time link to a browser booth, and record or re-record individual lines to replace AI delivery where it matters. |
| Score it with music | Generate original, royalty-free instrumental music on demand, or choose from a built-in library. Place intro, outro, transition, and interlude cues with precise fades and overlaps - music automatically ducks under speech so dialogue stays clear. |
| Mix and master automatically | Speech, recorded takes, and music are assembled into one timeline and finished with studio processing - reverb, compression, ducking, and broadcast-style loudness leveling - then exported as WAV or M4A. |
| Catch its own mistakes | After rendering, Autocast transcribes its own audio and checks it against the script, then automatically re-renders any dropped, repeated, or garbled lines - catching glitches a human editor would otherwise have to hunt for. |
| Finish and publish | Generate cover art and edit the title and description, preview a single section or the full episode, save every version as a draft, and publish straight to your uStudio media library - with re-publish and revert. |
How it works: a high-level view
Autocast moves an episode through a clear, mostly automated pipeline. You can step in at any stage to edit, direct, or record.
- Ingest - your documents, URLs, or prompt are read and understood, and the key points, themes, and narrative are extracted.
- Draft - an outline is planned (title, summary, hosts, chapters), then full multi-host dialogue is written for each chapter, grounded in your source.
- Direct & edit - refine the screenplay by hand, swap out hosts, add per-line performance cues, or describe changes in plain English.
- Voice & record - AI voices render each host; real people and live AI-host interviews can be recorded and folded in.
- Score - original or library music is generated and placed, ducking automatically under speech.
- Mix, check & publish - everything is assembled, mixed, and mastered; a self-healing quality pass repairs any glitches; then you preview, generate cover art, and publish to uStudio.
The models behind it
Autocast is built end-to-end on Google’s generative media stack within uStudio’s secure cloud infrastructure. The models below work together across the pipeline:
| Model | Role in Autocast |
|---|---|
| Gemini | Reads and analyzes your source material, generates the multi-host script, and renders the multi-speaker voices. |
| Gemini Live (native audio) | Powers the real-time conversations in which an AI host interviews a live human guest. |
| Lyria | Generates the original, vocal-free instrumental music. |
| Imagen / Gemini image | Generates episode cover art. |
| Speech-to-Text | Drives the self-healing audio quality check and live captions during recording. |
This initiative is part of uStudio Labs and is currently classified as a Concept Car.
Concept Cars are early-stage prototypes designed to help us explore new ideas, workflows, and opportunities in collaboration with customers. The goal is not to perfect a solution, but to learn what problems are worth solving and which approaches create the most value.
As with all uStudio Labs initiatives, customer feedback plays an important role in shaping what comes next.


