Case study · M.Tech · Generative AI

Text to video generator project

This text to video generator project is a delivered M.Tech generative AI system: you type a topic, choose a genre, tone and number of scenes, and it returns a narrated 1080p MP4.

Last updated

IEEE-format results page from the text-to-video generator project: output tables, mean opinion scores, a per-stage latency chart and a latency-scaling plot.
IEEE-format paper
Six AI-generated story frames from the text-to-video pipeline: a dinosaur skeleton in a museum hall, a jungle waterfall, an astronaut on Mars, a rocket launch, a surreal canyon between two giant brains, and a creature hatching from an egg.
Generated output
Story prompt screen of the AI video storyteller app: a topic box, genre and tone menus, a scene-count slider and a Generate Video Story button.
Web app

From this delivered M.Tech project, shown with names and institute details removed. Select an image to zoom.

An LLM writes the scene-by-scene script, a diffusion (text-to-image) model draws each scene, neural text-to-speech reads it aloud, and Remotion with FFmpeg assembles the video, with review steps in between. The student received the web app and code, an IEEE-format paper, journal- and Springer-format versions, a thesis and a viva guide.

What is this project, at a glance?

The AI video storyteller in six lines
ItemDetails
LevelM.Tech
DomainGenerative AI, multimodal content generation
Problem typeText-to-video generation: one topic in, a narrated video out
Core methodsLLM script writing with validated JSON output; diffusion (text-to-image) scenes kept consistent with an anchor image; neural TTS with a fallback chain; timing computed from measured audio; pan-and-zoom motion and music
StackNext.js, React and TypeScript; Zod validation; REST API routes per stage; Remotion and FFmpeg rendering; swappable LLM, image and TTS providers; Python for the evaluation and fine-tuning scripts
DeliverablesWeb app and source code, IEEE-format paper, journal- and Springer-format papers, M.Tech thesis, technical report, viva guide, sample videos

What problem does this AI video generator project solve?

Making a short story or explainer video normally means writing a script, finding pictures, recording a voice-over and editing it all to time. Text to video generation using AI can skip those steps with end-to-end video models, but they give you little control over the script and cost far more per second of footage.

This project automates the whole chain while keeping a person in charge at each step. Three engineering problems shaped the design:

  • Visual consistency. When each scene image is generated separately, the main character changes face and clothes from scene to scene. A text prompt alone does not carry identity across generations.
  • Audio-visual sync. Each scene must stay on screen exactly as long as its narration, or sentences get cut off.
  • Provider reliability. The pipeline depends on outside AI services, and any of them can rate-limit, run out of credit or retire a model.

How does the pipeline turn a prompt into a video?

The system is a chain of stages. Each stage is its own API route, so it can be retried alone, and the user reviews the output before the next, more expensive step runs.

  1. Script. An LLM writes a title and, for every scene, one or two lines of narration plus an image description, returned as JSON and validated before use.
  2. Review. The user reads the script and edits any line before images are generated.
  3. Consistency refinement. Scene descriptions become shot briefs, and a fixed description of the cast, wardrobe and setting is added to every image prompt.
  4. Scenes. The first scene is generated as an anchor; the others are generated in parallel with that image attached as a reference. Any scene can be regenerated.
  5. Narration. Text-to-speech voices each line through a chain of providers and records the real length of each clip.
  6. Timing and render. Frame counts come from the measured audio. Remotion renders the scenes with pan-and-zoom motion and music, and FFmpeg encodes the MP4.
Text-to-video pipeline flowchart: user prompt, LLM narrative with an approve-or-edit loop, cross-modal consistency refinement, diffusion synthesis with an approve-or-regenerate loop, neural TTS, then composition with Ken Burns motion and music into a 1920 by 1080 MP4.
Pipeline flowchartEnd-to-end pipeline: user prompt, LLM narrative, consistency refinement, diffusion scenes, neural TTS and 1080p MP4 composition, with user review gates.

The version described in the paper used a Llama 3.3 70B model for the script and a Gemini image model for the scenes, as the flowchart shows. Every provider sits behind a shared adapter interface, so the code can switch the LLM, image or voice service without touching the pipeline, and a mock mode runs the whole flow with no API calls.

Architecture of the AI video storyteller in four stacked layers: Next.js React UI, a RESTful API and pipeline orchestrator, LLM, diffusion and TTS adapters, and Remotion plus FFmpeg producing an H.264 MP4.
ArchitectureFour-layer system architecture: Next.js UI, REST API orchestrator, LLM/diffusion/TTS adapters, and Remotion + FFmpeg composition.

What did we build?

The web app runs the pipeline as a short wizard. It starts with one form: a topic, a genre, a tone and a scene count, with narration available in English, Hindi and Marathi.

Create Your Story screen of the Video Storyteller app, with a story topic box, genre and tone dropdowns, a scene-count slider and a Generate Video Story button.
Web appApp screen for creating a story: enter a topic, pick genre and tone, set the scene count and generate a narrated video.
  • Review screens for the script and for every generated scene, with inline editing and one-click regeneration.
  • An in-browser preview player, a server-side render with progress updates, and an MP4 download.
  • Fallback chains for images and speech that remember providers which fail permanently (a bad key, no credit) but never drop a provider over a temporary rate limit.
  • A label on every image naming the provider that made it, and a strict mode that fails loudly instead of quietly using a lower-quality free fallback.
Grid of six AI-generated frames: a museum dinosaur skeleton, a jungle waterfall, an astronaut on Mars, a rocket launch, a surreal brain canyon and a hatching creature.
Generated outputAI-generated story frames from the pipeline’s own outputs, covering museum, nature, Mars, rocket launch, surreal and creature scenes.

How was it evaluated, and what did the results show?

Evaluation had two parts. Human raters scored generated videos across eight genres on narrative coherence, visual consistency, audio quality, sync accuracy and overall production, reported as mean opinion scores. The pipeline was also timed stage by stage and at different scene counts.

  • Sync accuracy rated highest, because scene length comes from the measured narration rather than an estimate.
  • Visual consistency rated lowest and varied most by genre: nature and documentary stories held together best, fairy tales and adventure stories least.
  • An ablation with and without the consistency-refinement step showed that the step clearly improved rated coherence and consistency.
  • Image generation was the slowest stage, and total time grew almost in a straight line with the number of scenes, so the time for a longer video is predictable.
Grouped bar chart of mean opinion scores for the generated videos’ narrative coherence and visual consistency across documentary, fairy tale, sci-fi, historical, nature, educational, mystery and adventure genres.
ResultsGenre-wise mean opinion scores for narrative coherence and visual consistency across eight story genres.
Two-column IEEE-format paper page with a video output configuration table, a per-genre output summary, a mean opinion score table, a per-stage latency bar chart and a latency-versus-scene-count plot.
IEEE-format paperIEEE-format results page with output tables, mean opinion scores, a per-stage latency chart and a latency-scaling plot.

The write-up is also honest about limits: pan-and-zoom over still images is not true camera motion, consistency is enforced against a single anchor frame, and image generation dominates both time and cost. The project also set up a small-model experiment, with a teacher-generated dataset and a LoRA fine-tuning notebook for a compact open model that could write scripts offline; the full training run is left as future work.

Shown with student, guide and institute details removed.

What did the student receive?

Files handed over for this project
DeliverableWhat it contains
Web app and source codeNext.js app, pipeline stages, provider adapters with fallbacks, Remotion video composition, mock mode for offline demos
IEEE-format paperTwo-column paper with the method, quality ratings, latency analysis and an ablation of the consistency step
Journal- and Springer-format papersThe work reformatted for a journal, and a focused paper on keeping characters consistent across scenes
M.Tech thesis and technical reportLaTeX thesis, plus a technical report and architecture notes covering design decisions and measured performance
Viva guidePlain-language walkthrough of every stage, key numbers and likely examiner questions with answers
Sample outputsRendered MP4 videos and generated scene images for the demo

Stage reports, a black book or slides can be added for your own project through our dissertation, report and PPT support.

How could you adapt this project to your own topic?

The same pipeline carries over to many generative AI M.Tech topics. These are starting points to discuss with your guide.

Topic ideas, not delivered projects

The ideas below are suggestions. Our delivered work is in the case studies.

  • Topic ideaRegional-language explainer videosShort lessons on school topics in Indian languages, with narration quality rated by native speakers.
  • Topic ideaAn automatic consistency scoreMeasure style and character identity across scenes with image and face embeddings, then check the score against human ratings.
  • Topic ideaReal motion instead of stillsAdd a short image-to-video model per scene and compare it with pan-and-zoom on cost, time and rated quality.
  • Topic ideaA small offline script writerFine-tune a compact open LLM with LoRA on teacher-written scripts and compare it with the hosted model.

A lighter version, with fewer stages and a single provider, can also be sized for B.Tech: see final year projects for CSE. For more M.Tech scope, see M.Tech projects.

Frequently asked questions

Can I get a similar text-to-video project for my M.Tech?

Yes. We scope a text to video generator project with you in the free consultation around your own topic, base paper and your guide’s requirements, so the result is your project rather than a copy of this one. The quote lists each deliverable: code, paper, thesis and viva preparation.

Which models and datasets does it use?

It calls pretrained, hosted models through APIs: an LLM for the script, a text-to-image model for the scenes and a text-to-speech service, each behind a swappable adapter. The core pipeline needs no training dataset; the evaluation used human ratings of generated videos and timing measurements.

Do I need a GPU or paid API keys?

No GPU is needed, because the heavy models run as hosted services and the app runs on a normal laptop. Good-quality image generation usually needs a paid API key that is billed per image; a mock mode runs the whole pipeline without any API calls for development and demos.

What will I need to explain in the viva?

The order of the stages and why narration must finish before timing, how the anchor image and the fixed cast description keep characters consistent, what happens when a provider fails, how Remotion turns React components into video frames, and how the rating study was designed. The student’s plain-language viva guide covers each of these points.

Can this project become a research paper?

This one did: it produced an IEEE-format paper plus journal-format and Springer-format versions, built around consistency across independently generated scenes. We prepare the paper and help you choose venues, but acceptance is decided by the journal or conference.

Delivered work

Related case studies

More M.Tech projects we delivered, shown with names and institute details removed.

Planning a generative AI project? Bring your idea.

Share your level, topic and deadline. You get a written plan and a fixed quote, and the consultation is free.

B.Tech · M.Tech · PhD · Projects · Papers · Thesis · Reports

WhatsApp Free consultation