Imagera Hermes Maxa finished clip in about seven seconds
Fast enough to change how you work: try six framings of a shot in the time one render used to take, keep the one that landed, then finish it elsewhere if you need 4K. Every spec, every credit price and copy-ready prompts for all three modes — on this page.
What is Imagera Hermes Max?
Quick Answer:
Imagera Hermes Max is a video model on Imagera built for speed: a 5-second clip with synchronized sound returns in roughly seven seconds, from 45 credits. It runs text, image and reference modes, 5–15 seconds, up to 768P. It does not go to 4K — that is what makes it the cheapest and fastest video in the catalog.
Imagera Hermes Max is a video generation model on Imagera built around speed: a 5-second clip with synchronized sound comes back in about seven seconds. It runs three modes — text to video, image to video, and reference to video — at 480P or 768P, and you can try all three in the Imagera Sandbox.
How fast is it really?
A 5-second clip at 768P returns in roughly seven seconds, against minutes for most video models. That is the reason to choose it: it makes iterating on a shot practical, so you can look at six versions instead of imagining five of them.
Does it generate sound?
Yes — dialogue, effects and ambience are written with the picture in the same pass rather than dubbed on afterwards, so a line to camera arrives already lip-synced. There is no toggle and no surcharge; every clip comes back with audio.
What resolutions does it support?
480P and 768P, defaulting to 768P. There is no 2K and no 4K on this model. That ceiling is deliberate and it is what makes it the fastest and cheapest video option here — when a take has to be delivered in 4K, Imagera Hermes Video Gen is the same family with the higher ceiling.
Imagera Hermes Max specs at a glance
Modes
Text to Video · Image to Video · Reference to Video
Speed
A 5-second clip with sound returns in about seven seconds. Speed is the point of this model — it is what the ceiling below buys you.
Duration
5, 8, 10, 15 seconds (default 5s). 5s is the floor — shorter is refused, not rounded up.
Resolutions
480P · 768P (default 768P). There is no 2K and no 4K on this model — for those, use Imagera Hermes Video Gen.
Aspect ratios
21:9, 16:9, 4:3, 1:1, 3:4, 9:16 in Text mode. Reference mode offers the same 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 plus an adaptive option that follows your references. Image mode has no picker — the uploaded frame dictates the geometry.
Audio
Native synchronized audio — dialogue, effects and ambience written with the picture in the same pass, never dubbed on after. There is no switch and no surcharge: every clip comes back with sound.
Prompt length
Up to 2,000 characters — a full shot list with timing and sound cues fits
Frames (Image mode)
One start frame required, plus an optional closing frame — 2 images max, 10 MB each. With both, the clip animates the whole arc between them.
References (Reference mode)
Up to 9 images, 3 video clips and 3 audio takes — 12 files total. Cite each one by number in the prompt. An audio reference needs at least one image or video reference alongside it.
Repeatability
A seed is available in Advanced settings. Reuse it with the same prompt to return to a take you liked instead of rerolling — which matters here precisely because runs are cheap and fast enough to iterate.
Prompt expansion
Two modes, balanced and quality, defaulting to balanced. Quality rewrites your prompt before generating and can spend longer doing that than the render itself takes — leave it on balanced unless you want that trade.
Content filter
On by default and toggleable per run, with a one-time confirmation the first time you turn it off.
What is Imagera Hermes Max? Three modes, one model
Text to Video
Imagera Hermes Max
A finished clip in about seven seconds
A five-second clip with sound comes back in about seven seconds, where the rest of the lineup takes minutes. Sound is written with the picture in the same pass, never dubbed on after. It tops out at 768P rather than climbing to 4K. Six aspect ratios from 21:9 to 9:16, five to fifteen seconds, and the same engine takes a photo, a first-and-last frame pair, or a reference stack — the stack Hermes Turbo does not have. Use Turbo to find the shot; use this when you need references; use Hermes when you need to finish it in 4K.
Image to Video
Imagera Hermes Max (Image)
Your still, moving, almost immediately
Drop in a photo and it moves — in about seven seconds, where most engines here take minutes. Add a second image and it becomes your closing frame, and the clip tweens between the two. Up to fifteen seconds at 768P with sound written in the same pass. Hermes Turbo is the cheaper first run; this is the one that still takes a reference stack on its sibling mode; Hermes is the one that goes to 4K.
Reference to Video
Imagera Hermes Max (Reference)
Compose from photos, clips and voice — fast
Nine reference photos, three reference clips and three voice takes in one shot, cited by number in the prompt so each reference binds to its role: the face from photo 2, the camera move from clip 1, the delivery from take 3. The face you upload is the face on screen from the first second to the last — and the answer comes back in seconds rather than minutes. That speed is the moat on this mode especially: reference work is where you normally burn a morning re-running one shot to get the identity right, and here ten attempts cost less time than one used to. 768P ceiling, from 45 credits; step up to Hermes when the take is chosen and you want it in 4K.
Pricing in credits
Every run
One table — this model is priced the same whichever delivery speed is selected.
Resolution
5s
8s
10s
15s
480P
45
70
90
130
768P
60
95
115
175
All prices are in credits and always land on multiples of 5. Price scales with duration and resolution: the cheapest run is 5s at 480P (45 credits), the longest is 15s at 768P (175 credits). Sound is included at no extra cost, and reference files add nothing. Note that this model is not cheaper per second than Imagera Hermes Video Gen at the same resolution — a 5s clip at 768P is 60 credits on either. It costs less because it stops at 768P, so the expensive rungs simply are not there to reach.
Credits come in packs — see credit packs. No subscription required.
The problem with waiting minutes for a draft
Most AI video takes minutes per clip. That single fact shapes how people use it: you write one careful prompt, submit it, go and do something else, and come back to a result that is either right or a wasted wait. Iteration is expensive, so you stop iterating — and the shot you ship is usually the first one that was merely acceptable.
Imagera Hermes Max is built to remove that pause. A 5-second clip with sound comes back in about seven seconds, so trying six framings costs 270 credits at 480P and less time than one render used to. You find the shot by looking at shots, not by imagining them.
The trade is honest and worth stating plainly: this model stops at 768P. It is the tool for finding and testing a take, and for the enormous number of clips that never need to be larger than that. When a chosen take has to be delivered in 4K, Imagera Hermes Video Gen is the same family with the higher ceiling.
Prompt library
Copy any of these into the Sandbox as a starting point — each one exercises something this model is specifically good at.
Text to Video
6 prompts
Describe one shot, and name the sound you want — it is generated with the picture, so asking for it costs nothing extra. Because a run comes back in seconds, write loosely first and tighten on the second or third pass.
A potter looks up from her wheel in a sunlit studio and speaks straight to camera: "This one took me nine tries." Warm window light, shelves of bisque-ware behind her, gentle wheel hum under the line.
Dialogue with lip sync and room tone, all written in the same pass.
A red kayak crosses a mirror-still fjord at dawn, mist sitting on the water. Slow tracking shot from the shore, then the camera lifts to a high wide as the paddle breaks the reflection. Water, paddle drips, distant birds.
Landscape with a camera move and a natural ambient bed.
Slow dolly down a rain-soaked alley stacked with neon signage. A figure crosses under a clear umbrella, light breaking apart in the puddles. Rain on plastic, distant traffic, no music.
Night-time photoreal, reflections and wet surfaces — a stress test for the format.
A stag lifts its head from the leaf litter and turns to face the lens through shafts of light in an autumn beech wood. Handheld long lens, shallow focus, forest ambience and one distant call.
Wildlife documentary register, with a specific action in a specific order.
A blacksmith brings a hammer down on orange-hot steel and throws a shower of sparks. Macro, hard shadows, each strike ringing in time with the picture.
Fast motion with audio that has to land on the frame, not near it.
Vertical 9:16: a weathered fisherman mends a net on a harbour wall at first light, salt-stiff jacket and cracked hands. Shot with the grain and available light of 35mm reportage. Gulls, rigging, water against stone.
Vertical framing plus a documentary look rather than a rendered one.
Image to Video
5 prompts
Upload a start frame and the clip keeps its look. Add a second image and it becomes the closing frame, with the model animating the whole way between the two — a job that otherwise needs two renders and an editor in the middle.
Animate the provided frame: the figure walks on under the umbrella, the rain keeps falling and the reflections keep breaking. Lock the camera. Keep the signage and the colour exactly as in the still.
Extends a still into a shot with everything else held still.
From this forest frame: the stag lowers its head back to the leaf litter. The light shafts and mist behind it stay exactly where the still put them. Nothing else moves.
Single-subject motion — the clearest way to see the model respecting a source frame.
Use the first image as the opening frame and the second as the closing frame. Move between them as one continuous push-in, the light shifting from overcast to golden across the clip.
The first-and-last-frame arc, which is what the second image slot is for.
Bring this harbour-wall portrait into motion: he pulls the net cord tight, tests the knot, then looks up along the wall. His face is unchanged from the frame it started as.
Identity held from a photograph while the body performs an action.
The dish in the photo steams gently as a hand enters frame and drizzles sauce in a spiral. Macro top-down, then a slow tilt to a hero angle. Sizzle and kitchen ambience underneath.
Product motion with a camera move and a matching sound bed.
Reference to Video
5 prompts
Attach up to 12 files and cite each by number so every reference binds to a role — the face from photo 2, the camera move from clip 1, the delivery from take 3. This is the mode where the speed matters most: identity work normally means re-running one shot all morning.
Photo 1 is our founder; photos 2 and 3 show the product front and side. She carries it across a sunlit studio, sets it on the bench and looks up to camera. Keep her face exactly as in photo 1 and the product proportions from photo 2. 16:9, warm daylight.
Face lock and product lock in a single generation.
Recreate the camera move from clip 1 — the slow rising crane — over the skyline in photo 1 at dusk. Match its pacing and grade. Distant traffic and wind underneath.
Borrow motion from footage, apply it to a new scene.
Photos 1 through 6 are the same mascot from different angles. Animate it doing a short victory dance in the office from photo 7, timed to the rhythm of audio 1. Colours and proportions identical in every frame.
Multi-angle character lock — up to 9 images fit, which is what makes this affordable.
Match the speaker in photo 1 to the voice in audio 1: she delivers the line to camera in a bright studio, lips synced to the take, gentle push-in. 16:9, soft key from camera left.
Pairing a voice take with a face — remember an audio reference needs an image or clip alongside it.
Built from the single reference of the potter: she carries a finished bowl across the studio to the drying shelf and steps back to look at it. Same face, same apron, an action the reference never showed.
One reference photo, a shot that was never filmed.
Getting the most out of Imagera Hermes Max
Work in passes, not in one careful prompt
The reason to use this model is that a run costs seconds. Six drafts at 480P come to 270 credits and take less time than one render on a slower engine. Write loosely, look, then tighten — the looking is the technique.
Draft at 480P, finish at 768P
A 5-second draft is 45 credits at 480P against 60 at 768P. Iterate cheaply, then re-run the keeper at the higher rung.
Save the seed when a take lands
A seed sits in Advanced settings. Note it on a run you liked and reuse it with the same prompt to come back to that take rather than rerolling and hoping. This family is one of the few here that can do it.
Ask for the sound explicitly
Audio is written with the picture in the same pass and costs nothing extra, but the model gives you what you name. Put dialogue in quotes, name the ambience, and say what should land on the beat.
Leave prompt expansion on balanced
The quality setting rewrites your prompt before generation starts and can take longer doing that than the clip takes to render — which spends this model's whole advantage on rephrasing something you already wrote.
Crop the still before uploading in Image mode
There is no aspect-ratio control in Image mode: the frame you upload dictates the output geometry. Frame it the way you want the video framed.
Add a closing frame for a guaranteed ending
Two images turn Image mode into a first-and-last-frame job, and the model animates the entire arc between them. Product reveals and before/after shots are the obvious uses.
Cite every reference by number
Say what each file contributes — "face from photo 1, motion from clip 1, delivery from take 1". A reference the prompt never mentions is wasted signal, and this mode charges the same either way.
Know the floor and the ceiling
5 seconds is the shortest run the model accepts — ask for less and it is refused rather than rounded up. 768P is the top: there is no 2K or 4K here, so plan the finish accordingly.
How it works
STEP 1
Open the Sandbox and pick the model
Open the Imagera Sandbox and select Imagera Hermes Max — or the Reference variant if you are bringing photos, clips or voice takes.
STEP 2
Write a shot and name the sound
Type a prompt, upload a start frame, or attach references and cite them by number. Say what should be audible; sound is generated with the picture.
STEP 3
Draft cheap, then finish
Run 5s at 480P for 45 credits while you iterate, then re-run the take you kept at 768P for 60.
Use cases
Pre-production
Finding the shot
Try six framings of an idea in the time one render normally takes, then commit to the one that worked.
Social
Social clips
Vertical and square are both in the ratio list, and a 5-second clip with sound is 60 credits.
Brand
Talking-to-camera lines
Dialogue, mouth and room tone are generated together in one pass — no separate dubbing or lip-sync step.
Brand
Recurring characters
Lock a face or mascot from up to 9 reference photos, cheaply enough to run ten attempts.
Creators
Stills brought to life
Turn a product shot or portrait into a clip that keeps its exact look, or tween between two frames.
Agency
Storyboards and animatics
Cheap, fast clips are the right resolution for a board — nobody needs an animatic in 4K.
Every fact and credit figure in this table is read from the same registry the studio uses. Both models share native synchronized audio.
Available via the Imagera API
Imagera Hermes Max is available through the Imagera API as Imagera Hermes Max. The three modes map to three endpoints — same credits as the studio, billed per second.
Imagera Hermes Max
POST /v1/queue/imagera-video-hermes-max
Fast text-to-video with native audio, 5–15s at up to 768P across six aspect ratios.
Imagera Hermes Max — From Image
POST /v1/queue/imagera-video-hermes-max-from-image
Fast animation of a still, or tweens between a start and end frame when a second image is supplied. Up to 768P.
Imagera Hermes Max — From References
POST /v1/queue/imagera-video-hermes-max-from-references
Fast composition from up to 9 reference images, 3 clips and 3 audio takes, cited by number in the prompt. Up to 768P.
Endpoint schemas, keys and quickstarts live in the developer docs.
IA
Imagera AI Team
Unified AI creation platform
Frequently Asked Questions
What is Imagera Hermes Max?+
Imagera Hermes Max is a video generation model on Imagera built around speed: a 5-second clip with synchronized sound comes back in about seven seconds. It runs three modes — text to video, image to video, and reference to video — at 480P or 768P, and you can try all three in the Imagera Sandbox.
How fast is it really?+
A 5-second clip at 768P returns in roughly seven seconds, against minutes for most video models. That is the reason to choose it: it makes iterating on a shot practical, so you can look at six versions instead of imagining five of them.
Does it generate sound?+
Yes — dialogue, effects and ambience are written with the picture in the same pass rather than dubbed on afterwards, so a line to camera arrives already lip-synced. There is no toggle and no surcharge; every clip comes back with audio.
What resolutions does it support?+
480P and 768P, defaulting to 768P. There is no 2K and no 4K on this model. That ceiling is deliberate and it is what makes it the fastest and cheapest video option here — when a take has to be delivered in 4K, Imagera Hermes Video Gen is the same family with the higher ceiling.
How much does it cost in credits?+
From 45 credits for a 5-second clip at 480P, up to 175 credits for the full 15 seconds at 768P. The default 5-second clip at 768P is 60 credits. Sound and reference files add nothing, and every price lands on a multiple of 5 credits.
Is it cheaper than Imagera Hermes Video Gen?+
Not per second at the same resolution — a 5-second clip at 768P costs 60 credits on this model and 60 on Imagera Hermes Video Gen. It is cheaper in practice because it stops at 768P, so the expensive 2K and 4K rungs are not there to reach. You are choosing a ceiling, not a discount.
How long can a clip be?+
5 to 15 seconds, with presets at 5, 8, 10, 15 seconds. 5 seconds is a hard floor — a shorter request is refused rather than rounded up — and price scales per second.
Can I set a first and last frame?+
Yes. Image mode takes a start frame and an optional closing frame (2 images max, 10 MB each) and animates the whole arc between them. With one image it simply animates from that frame.
How many reference files can I use?+
Up to 9 images, 3 video clips and 3 audio takes, capped at 12 files in total. Cite each one by number in the prompt and say what it contributes. An audio reference cannot be the only reference — it needs an image or a clip alongside it.
Can I get the same result twice?+
Yes — a seed is available in Advanced settings. Reuse it with the same prompt to return to a take you liked rather than rerolling. That pairs well with how fast runs are: you can afford to explore, then go back.
Which aspect ratios are available?+
Text mode offers 21:9, 16:9, 4:3, 1:1, 3:4, 9:16. Reference mode offers the same set plus an adaptive option that follows your references. Image mode has no picker — the frame you upload sets the geometry, so crop it first.
What is prompt expansion?+
A setting with two modes, balanced and quality, defaulting to balanced. The quality mode rewrites your prompt before generating, which can take longer than the render itself — so it trades away the speed you came for. Leave it on balanced unless you specifically want that.
Can I use it through the API?+
Yes — as Imagera Hermes Max (imagera-video-hermes-max), with sibling endpoints for the image and reference modes. Same credit prices as the studio. See the developer docs at imagera.ai/developers.
Try Imagera Hermes Max now
Pick it from the model rail, paste a prompt from this page, and watch one continuous take come back with its own soundtrack.