Veo 3.1 is a video generation model from Google. Its standout ability is sound: it creates lifelike clips where speech, ambience and sound effects come out of the same generation as the images, so lips, footsteps and background noise line up with what happens on screen. On NeoMundo, Google Veo sits next to 12 other image and video models in one studio and one credit balance.
The model accepts a text prompt, which is required, plus up to 2 reference images. With no images it works as text to video. With one image it animates your picture from that opening frame. With two images it uses start and end frames, which lets you decide exactly where the shot begins and where it lands. It does not take video or audio uploads; everything you hear is generated from your prompt.
Output options are deliberately compact: 16:9 or 9:16, 720p, 1080p or 4k, and a duration of 4, 6 or 8 seconds. Veo 3.1 starts at 66 credits for the smallest combination, and the default setup costs 132 credits. Longer clips and higher resolutions cost more, and the studio shows the exact figure before you start.
Because each clip is short, the model works best when you think in shots rather than whole stories. Generate one moment at a time, keep the ones you like, and cut them together in your editor. Several shots can run in parallel, since each account can have up to 8 generations running at once.