Houdini

From Houdini to Gen AI: How I Rendered a Film-Grade USS Missouri Ocean Shot in Under 24 Hours

We’ve all been there. A brief comes in for a film or TV project requiring a massive, high-detail ocean shot—in this case, the legendary USS Missouri (BB-63) cutting through the water.

Normally, this is the cue for your workstation to melt. You’re looking at heavy FLIP fluid simulations for the ocean, secondary whitewater/foam simulations, and Pyro sims for the diesel smoke chugging out of the funnels. It’s a process that can easily eat up days of setup, caching, and rendering.

For this project, I decided to test a different hypothesis: Can Generative AI replace heavy Houdini simulations for quick, production-ready rushes while maintaining absolute camera and animation control?

The short answer is yes. But the secret isn’t just writing a clever prompt, it’s how you prep your 3D assets to act as the ultimate control rig. Here is how I went from a raw 3D model to a finished, color-graded shot in just one day.

1. The Strategy: 3D as the Ultimate Guide

In film and TV, camera direction and object animation are non-negotiable. “Garbage in, garbage out” is the golden rule of AI. If you feed an AI video generator a loose, low-res viewport flipbook, the AI has to guess too much, resulting in warping and boiling.

I needed a rock-solid reference video. I decided to frame the shot as a high-altitude aerial photo-op, the kind navies use to showcase their fleets (a composition I actually got to know intimately during my own navy days).

Example photo ops of naval ships (Credit: Military Tech YouTube Channel)

To give the AI the best possible guide, I opted for a semi-full Karma render in Houdini (with lighting and textures, but absolutely zero simulation elements).

The Houdini Prep & The “Reflection Grid” Trick

I loaded up the USS Missouri model, adapted its OBJ textures for Houdini’s Karma renderer, and animated a sweeping aerial camera path. I also added some subtle, realistic movement, animating the A and B gun turrets rotating slightly as the ship cruised forward.

However, rendering a ship on black space doesn’t work for AI. I needed to establish lighting and reflections:

  • The Reflection Problem: I lit the scene to look like late afternoon with warm, golden-hour sunlight. But without water, the ship’s metallic hull looked flat. I placed a ground grid directly beneath the ship and applied a material that mimicked the reflective properties of ocean water. Immediately, the warm sun bounced beautifully off the hull.
  • The Motion Problem: With a featureless reflective grid, the moving ship looked completely stationary as the camera tracked it. To give the camera (and the AI) a sense of speed, I needed a visual reference.
  • The Solution: I duplicated the grid, turned it into a wireframe white grid, and overlaid it. To prevent these white grid lines from ruining my beautiful hull reflections, I adjusted the geometry render settings so the wireframe grid was only visible to primary rays.

The result? Perfect hull reflections, a clear grid to register relative motion, and a clean 1080p render sequence ready for the AI pipeline.

The semi-full render from Houdini to be used as reference video later

2. Setting the Stage: The First Frame (FF)

Because the first frame dictates the visual quality of the entire video, I needed to replace our placeholder grid with actual, high-fidelity ocean waves, foam, and smoke.

Using my rendered first frame as an image prompt, I wrote a highly specific control prompt to dictate the physics of the scene:

“Change grid floor to ocean. Give the bow wave and wake solid, clearly visible whitewater — a defined white bow wave breaking at the hull with decent foam volume, and a wake trailing behind the ship with an appropriate amount of bright, clearly visible whitewater, picking up warm tones from the light where it catches the foam. Keep this consistent with a heavy displacement-hull warship at moderate cruising speed, not a high-speed planing wake. Foam edges should be organic and irregular, breaking up naturally further behind the ship.

Add two separate columns of dark grey-black smoke, one rising independently from each of the two black funnels, both blown by the same wind, streaming backward at the same consistent angle, never merging into one plume, with clear open space between them — the smoke’s edges may pick up a subtle warm rim from the low sun. Follow the exact light direction and warmth already present in the source image.”

The first output from Nano Banana Pro

Iteration & Tweaking

The first output generated a wake that was way too heavy—it looked like the battleship was trying to drag race. I quickly re-prompted to scale it back:

“Make the smoke columns trail longer towards the back of the ship. 20% less white water.”

The improved output from Nano Banana

This hit the absolute sweet spot.

3. Bringing It to Life in Kling Omni

With my master assets ready—the perfect First Frame (FF) and my 1080p Houdini reference video—I loaded them into Kling Omni.

I ran the generation using the reference video for spatial/motion information and the image as the exact starting frame, using the prompt:

“Use @Image as exact first frame. Use @Video as reference video and spatial information. The bow wave and wake should have solid, clearly visible whitewater.”

The Output

Output from Kling

The results were incredibly impressive. The water wake looked highly realistic, the lighting remained perfectly consistent as the camera moved, and the dual columns of funnel smoke drifted naturally in the wind.

Fixing the Details in Post

As expected with current Gen AI, fine details like the “63” hull numbering and the American flag printed on the turret had some slight consistency issues over time. However, because Kling followed my 3D reference video so precisely, fixing this was a breeze.

I took the original high-res clean render from Houdini and simply composited the stable decals, flag, and hull numbers right back over the AI video in post. Add a quick color grade, and the shot was complete.

The final result after color grading and compositing

Bonus Test: What Happens When You Push the First Frame Parameters?

To really test how much influence the First Frame (FF) has over the video engine, I ran a quick experiment. I took the earlier iteration of our first frame—the one with the much heavier, highly aggressive bow wave—and ran it through Kling Omni using the exact same prompt and 3D reference video.

The result was a textbook demonstration of how AI interprets initial input data:

  • 1:1 Motion & Physics Translation: The AI immediately locked onto the heavier wake from frame one and maintained that high-energy, violent water displacement across the entire camera move.
  • Zero Loss in Camera Alignment: Despite the drastic change in water volume and foam density, the model followed the 3D reference video’s spatial motion just as flawlessly as the calmer version.

The Resulting Shot

Here is the high-energy output:

Bonus footage of the output that was produced using a first frame image with bigger and stronger wake

Conclusion: The “First Frame” Controller Paradigm

This project completely shifted how I view the integration of 3D and Generative AI. If you need to deliver high-quality, complex shots on a tight deadline, here are my biggest takeaways from this experiment:

  • A Film-Grade Alternative for Quick Rushes: For elements that typically demand heavy, time-consuming simulations in Houdini, Gen AI is a massive time-saver. Going from a raw 3D model to a finished, client-ready shot in under 24 hours is a massive win for tight production schedules.
  • The First Frame is Your “Houdini Parameter Slider”: The most crucial realization was that when you pair a First Frame (FF) with a reference video, the FF acts as your physical controller for the “simulation.” If your first frame has high wake and heavy foam, the AI maintains that high-intensity motion throughout the shot. If you want a calmer sea, you dial it back in the first frame. Your control over that initial image is everything.
  • Keep Your Video Prompts Simple: When it’s time to prompt Kling, less is more. You don’t need to write a massive block of text describing the scene’s physics. Because you used a high-quality reference video and a detailed first frame, the spatial information, motion, and visual details are already completely baked into your inputs. The AI just needs a simple nudge to do its job.
  • The Power of the Semi-Full Render: This was my first time testing a semi-full render (complete with lighting, textures, and reflection grids, minus the actual simulations) as an AI input, rather than a flat viewport flipbook. The experiment proved that feeding the AI proper lighting and surface reflection data is the key to unlocking true, production-grade output.

By letting Houdini handle the rigid composition, perspective, and lighting, and letting Gen AI handle the chaotic, computationally expensive fluid dynamics, you get the absolute best of both worlds.

Continue Reading

Can AI Handle True Physics? Testing Refraction & Dynamic Shadows (Houdini vs. Kling 3.0 vs. Seedance 2.0)

We’ve all seen the impressive AI video clips floating around lately. By combining a First Frame (FF) image with a Reference Video, creators are getting incredibly believable outputs, especially when the reference footage relies on simple, solid geometries.

But believability isn’t the same thing as physical accuracy.

As AI models get better at mimicking motion, a major question remains: How much can we trust AI to properly handle physics-based rendering, like complex light transmission, refraction, and dynamic shadows?

To find out, I set up a strictly controlled environment to test whether modern AI video generators actually “understand” light, or if they just hallucinate pretty pixels.

The Setup: Establishing the Ground Truth

To avoid any ambiguity, I bypassed real-world camera footage entirely. Instead, I built a control test completely in Houdini, complete with dialed-in materials, physically accurate lighting, and a moving light source.

The Scene: A bright, spherical light source translating horizontally behind a transparent orange cube, casting dynamic soft, colored shadows. The Inputs for AI:

  • Reference Video: The raw Houdini Flipbook viewport preview (to provide the exact geometric motion path).
  • First Frame: The first frame of the fully rendered and denoised Houdini sequence.
  • Prompt: “Moving bright ball of light. Realistic shadow cast by the transparent orange cube. Use reference image as first frame.”
Houdini Flipbook render
First Frame image used to prompt the Gen AI models

I ran these identical inputs through two of the leading multi-modal video generators: Kling 3.0 and Seedance 2.0. Here is the breakdown of the results.

The Results Analysis

1. The Ground Truth (Houdini + Denoise)

Before judging the AI, let’s look at how actual physics behaves. In the Houdini render, as the bright sphere moves from left to right behind the transparent cube, the resulting colored shadow shifts its angle in exact, inverse correlation to the light source. The orange tinted shadow projected onto the ground is physically accurate and scales logically as the light source changes position.

Houdini control render shot

2. Kling 3.0: Aesthetically Pleasing, Physically Almost There

Kling 3.0 output

Kling 3.0 does a highly commendable job of maintaining the initial aesthetic established by the first frame. The texture of the ground and the glass-like quality of the cube remain quite stable, and the overall movement of the ball and the reflections off the cube are remarkably similar to the control footage.

The Physics Breakdown: Kling actually does a surprisingly decent job mimicking shadow-casting logic for a majority of the clip. It manages to infer depth; when the ball travels further to the right of the frame, Kling correctly elongates the shadow, matching the behavior of the Houdini control footage. However, its understanding of absolute spatial coordinates breaks down at specific points in time. As seen in the reference frame below, when the ball is directly behind the cube, Kling fails the geometry test and skews the shadow to the left, rather than casting it directly in front of the cube. It is approximating the physics based on visual context, rather than calculating actual spatial relationships.

Kling skews the shadow to the left despite the light source directly behind the cube.

3. Seedance 2.0: Loose Tracking and Creative Liberties

Seedance 2.0 output

Seedance 2.0 seems to have a mind and opinion of its own. Despite being fed the same prompts and reference video, it actively decided to ignore the 1:1 motion constraints and initiated a camera pan instead.

The Physics Breakdown: While it abandoned the tracking data (the reference video clearly shows the relative size and position of the ball relative to the cube, yet Seedance interprets the ball’s movement as perfectly symmetrical from left to right), the execution of its own idea is impressive. It handles the camera pan smoothly, and the lighting and material rendering look completely comparable to both Kling and the Houdini control.

The Verdict

The true state of AI video generation isn’t that it completely fails at physics, but rather that it relies on highly educated, aesthetic approximation instead of mathematical simulation.

Kling 3.0 proves that AI can infer complex physical relationships, like elongating a shadow as a light source moves further away, but it will still slip up on basic spatial alignment when it lacks a true 3D coordinate system. Seedance 2.0 proves that these models can render beautiful, accurate materials and lighting, but they often prioritize cinematic flair over strict adherence to your control inputs.

Ultimately, if you do not need exact 1:1 motion tracking with your control footage, tools like Seedance successfully fail the mission by providing gorgeous, usable alternatives. But if you are working in a pipeline that demands absolute precision, where a light must accurately cast a shadow directly beneath it, AI is still a black box of creative liberties. For true control, procedural workflows remain undefeated.

Continue Reading

AI vs. Physics: Can Video-to-Video Models Handle Fluid Refraction and Caustics?

As a visual effects artist, procedural fluid simulations are a staple in my workflow. Building custom velocity fields, tweaking flip solvers in Houdini, and dialing in physically accurate rendering for water takes time, computational power, and a lot of patience. With the rapid explosion of AI video-to-video models, a natural question arises: can these tools bypass the heavy lifting of traditional fluid sims?

Specifically, I wanted to test whether AI can accurately interpret water movement, the refraction of light through a volume, and water caustics. To find out, I set up a stress test comparing a traditional procedural workflow against two AI models: Seedance 2.0 and Kling 3.0.

Here is a breakdown of the experiment and the results.

The Control: Procedural Accuracy in Houdini

To establish a baseline, I built a controlled simulation in Houdini. The setup was straightforward but designed to test specific optical properties:

  • A small ocean geometry generating surface waves, placed over a simple grid.
  • A yellow cube placed on the grid, fully submerged under the water.
  • Materials accurately assigned to calculate the Index of Refraction (IOR).
  • Caustics explicitly turned on and lighting set to highlight the light patterns on the floor.
Houdini Setup to make the Control Shot

The scene was rendered in Karma. The very first frame of this render served as the ground truth (and also reference image) for how light should bend around the submerged cube and cast caustics on the ocean floor. The whole 3-seconds shot was also rendered out and composed for comparison with other Gen AI outputs.

The final shot that was made using Houdini and composed in After Effects.

The Methodology

To test the AI, I provided the models with the exact same starting point and motion data.

  1. The Input Frame: The extracted first frame from the Houdini render.
  2. The Reference Motion: An extracted flipbook animation (mp4) of the waves and the motion of the small ocean geometry.
  3. The Prompt: Use ff as first frame. ff shows a yellow cube in shallow water with water caustics, distorted by refraction of light. Use Video1 as reference video for movement of water and position of cube. Interpret the distortion of the cube and water caustics. No audio.
The Flipbook render from Houdini used as reference video

This exact combination was fed into both Seedance 2.0 and Kling 3.0.

The Results: How Did the AI Perform?

Evaluating the outputs required looking past the initial “wow” factor of AI generation and critically analyzing the physics of the scene.

Seedance 2.0: Good Motion, Broken Physics

Seedance 2.0 output
  • Water Movement: Seedance did a surprisingly good job imitating the motion of the waves from the flipbook reference. It definitely looked like a fluid, and the pure white reflections scattered across the surface were convincing.
  • Material Properties: This is where the physics fell apart. Instead of rendering clear water in a pool, Seedance generated what looked like an opaque, sky-blue liquid.
  • Refraction: Because the model hallucinated a highly opaque liquid, the physics of light transport were completely illogical. A thick, opaque liquid shouldn’t allow light rays from a submerged cube to refract that clearly. Setting that aside, the distortion on the cube was extremely minimal compared to the Karma render. It appeared the model simply morphed the object toward the yellow square I left in the reference video for general positioning. (Note for future tests: removing the tracking square might give the AI more freedom to interpret genuine optical distortion rather than acting as a rigid morph target).
  • Water Caustics: Due to the opaque nature of the generated fluid, the caustics were almost entirely non-existent.

Kling 3.0: Total Hallucination

Kling 3.0 output
  • Output: The result from Kling 3.0 was completely unusable for this specific use case. The model entirely misinterpreted the prompt and reference, generating what looked like a frozen layer of ice with water running underneath it. While there were some vague caustic-like patterns on the pool floor and distortion of the cube, the overall output was a failed interpretation of the input data.

Conclusion

While video-to-video AI models are making incredible strides in tracking 2D surface displacement, this test highlights a massive gap in their ability to understand volumetric data and complex physical light transport.

The AI tools treated the reference as a 2D warping task. Seedance 2.0 successfully mimicked the surface fluidity but failed to comprehend the depth of the water, fundamentally breaking the material properties by turning a clear, refractive medium into an opaque liquid. Kling 3.0 simply hallucinated entirely different physical states.

Feature TestedSeedance 2.0Kling 3.0
Water MovementConvincing surface tracking and fluid motion.Unusable (hallucinated an ice layer).
Material AccuracyFailed (rendered an opaque liquid).Failed.
RefractionMinimal; hindered by opacity and reference morphing.Some distortion present but unnatural-looking.
CausticsBarely visible due to incorrect material density.Failed.

For now, if a project requires precise optical fidelity, where refraction and caustics need to accurately interact with submerged objects, it still absolutely demands a robust procedural solver and a physically based renderer. AI can approximate the broad strokes of movement, but true fluid physics remains securely in the realm of traditional 3D pipelines.

Continue Reading