Product thinking

Prompts Step Back, Agents Move Forward

Agent mode turns a goal, template, or example into a guided video project, while the Drama workbench Agent helps creators control scripts, assets, storyboards, and production.

AI-generated short videos are becoming part of daily content streams.

They may not all be polished, but the formula is direct: a strong premise, a character image, and a few lines of dialogue can become a compelling short video.

For creators, the threshold and cost of generating an interesting clip are falling.

Making one clip easier does not make an entire project easy. A single standout moment can come from inspiration. A complete series needs continuity across characters, locations, and shots, supported by a coherent world and stable visual assets.

So we want to throw out a judgment first:

AI video generation will keep getting cheaper. Project context is what remains valuable.

That belief is where Vitent begins.

Vitent is an AI-driven engine for continuous content production. Creators can start with an idea, outline, or episode script and develop it into videos with stable characters and locations, a coherent visual style, and continuity from shot to shot.

Vitent now exposes agentic creation at two levels. Agent mode on the home page guides a complete video from an outcome, template, or Inspiration case. Inside a Drama project, the workbench Agent helps operate the script, asset, storyboard, and Production tools.

An Agent should really know what is happening in the project

We ran an English-interface workflow test in Vitent using a small fantasy project called Heavenly Corner Store. The captured test covered project setup, script editing, asset extraction, prompt review, four asset-image generations, storyboard generation, Production import, and one completed video clip.

The Agent first confirmed the project settings. After the script was added, it extracted Milo, Nim, the Heavenly Corner Store, and the Silver Lantern into the project asset library. We reviewed original, brand-neutral English prompts and generated one reference image for each asset. The same named assets were then linked to one scene and four storyboard clips from the final screenplay.

The captured workflow used direct instructions to extract script assets and create the storyboard. In Production, Import All carried the current storyboard prompt and three linked references into the video node. The 13-second, 9:16, 480p preview completed successfully and was accepted as clip 1.1 of four.

Reference the audited English script and review four generated asset images
Reference the English script, extract the reusable assets, and review the four saved image references.
Review the audited English Storyboard across four completed clips
Review the regenerated English scene, its four clips, and the project assets linked to each shot.
Import the English storyboard and references, configure a low-cost preview, and review the completed Production clip
Import the storyboard and its three references, confirm the settings, and review the completed 13-second preview.

The same named characters, location, and prop moved from the script into the asset library, storyboard references, and Production canvas without being re-entered in another tool.

The completed preview provides evidence for one end-to-end clip. Milo, the store, and the Silver Lantern remain broadly recognizable, but the lantern's light state is less consistent in the middle of the clip and generated lettering is not reliable enough for exact on-screen copy. Precise product names and legal text should be added as editor layers.

The captured evidence supports a precise conclusion: Vitent can carry named project context from a screenplay through reusable assets and a storyboard into a completed Production clip in one workspace. Full-episode and multi-episode consistency still require more completed clips.

Agent mode applies the same principle to guided creation

We also remixed the Inspiration case A Quiet Skincare Ritual in the English Agent interface. The Agent preserved a 15-second commercial structure, requested replacement inputs, generated a fictional serum and adult demonstrator, prepared the script and storyboard prompts, and delivered a completed 16:9 marketing video.

Remix an Inspiration case, approve replacement assets and a video plan, and review the completed Agent video
Remix a case structure with new demonstration assets and move through reviewable approvals to a finished video.

The LUMIÈRE product shown in this test is a fictional demonstration asset. A real campaign should replace it with owned product imagery, verified benefits, approved brand copy, and any required disclosures.

Video creation moves from “writing prompts” to “operating projects”

There are many video models today, and their capabilities are evolving rapidly. Of course, creators need to select a model and adjust the generation parameters. However, if the product only stops at "selecting a model, entering prompts, and waiting for the results," AI video creation will quickly fall into a new type of manual labor.

We feel there is a judgment worth debating here:

As AI video models improve, prompt tricks become less valuable while project context becomes more valuable.

The reason is simple. Better models reduce the effort required for a single generation. A shot that needs a complex prompt today may need far less description tomorrow. A scene that takes many retries today may work on the first attempt later. Model progress turns specialized prompt techniques into default capabilities.

But the project context does not appear automatically. A project's worldview, asset versions, shot cadences, sound styles, and multi-episode continuity all need to be organized, documented, referenced, and executed. The more complex this part is, the greater the value of the Agent.

We therefore see Vitent as the workspace for an AI video project. Models provide the engine, Agent mode guides outcome-focused creation, and the Drama workbench Agent operates inside structured production tools.

Vitent connects an Agent-led starting point with an Agent-assisted production workspace.

In Vitent, the Agent can read and update project scripts, initialize and revise assets, help design storyboards, and operate on the canvas.

When a creator says “continue with the next clip,” the Agent knows where the next storyboard is. When they say “keep this character unchanged,” the system knows which asset defines that character. When they ask for a scene to feel more oppressive, the Agent can locate the relevant location, characters, and storyboard before preparing the next generation task.

The creator is working with a project that has memory, structure, and production state. The AI can therefore behave more like an active project collaborator.

That may become the defining difference between AI video products.

Every supporting tool helps the Agent work better

The infinite canvas lays out an episode in one connected space. Storyboards, character images, location references, prop images, and video nodes can all be linked. Creators see the production workspace; the Agent reads structured context.

Asset management keeps characters, locations, and props in the project's long-term memory. They can be named, reused, and revised instead of disappearing into isolated generation results. This becomes increasingly valuable as a series grows.

The 3D virtual studio addresses spatial decisions before generation: where characters stand, where the camera looks from, and how the scene is arranged. The long-term opportunity is for the Agent to create an adjustable first pass from the script, storyboard, and directing language.

For example, a storyboard may specify an over-the-shoulder shot, shot-reverse-shot coverage, the relative position of two characters, and a camera move. The Agent should turn that information into adjustable blocking, camera placement, shot size, and composition. The creator can then refine the setup manually or continue through conversation. This preserves directorial control while reducing mechanical placement and repeated trial and error.

The goal of these capabilities is straightforward: let the Agent read the project's structure, retrieve the right assets, understand each shot, and move the work forward.

Vitent's first public note

Vitent is still early. The Agent needs smarter next-step suggestions, clearer task states, and stronger analysis and revision after generation. The 3D virtual studio is also at an early stage and will connect more deeply with storyboards, assets, and video generation.

But our direction is clear:

AI video will not stop at “enter a sentence and generate a clip.” The more meaningful change will happen at the project level.

Models will keep reducing the cost of individual generations. The larger advantage will come from continuous production: retaining project memory, organizing context, and coordinating work across assets, shots, and episodes.

This is what Vitent is building: an Agent inside the video project that moves creation from isolated generations to continuous production.

We are refining this workflow with individual creators, small businesses, brand and product marketers, AI animated-comic teams, and education content teams. We will continue to share product progress, real production processes, and practical AI video methods.

If you are exploring how AI video can move from one-off generation to continuous production, we would like to hear from you.

Try Vitent:https://app.vitent.ai/

This is our first public note.