A week ago Saturday, an invite-only group of executives, lawyers, agents, producers and technologists gathered at AFI in Los Angeles to discuss the state of AI. The conversation was held under the Chatham House Rule. No cell phones. Ben Affleck was the star attraction.
I wasn’t invited, and neither were you.
That has become a familiar story. Producers and filmmakers are constantly told that AI will disrupt every part of production, but almost nobody can explain what that means in practice.
During Netflix’s second-quarter earnings call, Ted Sarandos said the company had used generative-AI workflows on roughly 300 titles, primarily in postproduction. Doing what? Removing wires? Building crowds? Altering performances? Replacing locations? Netflix offered a few examples, but most of the work remains a black box. My guess is that a lot of it is still dubbing and subtitling.
Placing the bet
There are evangelists who believe generative AI will eventually absorb the entire production process. Then there are filmmakers betting on assisted AI: keeping actors, directors and conventional photography at the center while using the technology to extend, alter or finish what they capture.
Both paths will succeed and fail in interesting and unpredictable ways. But storytellers planning for the next few years cannot wait for the argument to resolve itself. We have to decide where to put our time, money and attention.
I have become increasingly interested in assisted-AI storytelling.
Neither path is harmless to the people who make movies. Fully generated work can replace jobs outright. Assisted AI can shrink crews, eliminate shooting days and move work away from departments that have supported generations of craftspeople. Those are not abstract efficiencies. They are jobs, health insurance and careers.
I don’t have a satisfying answer to that.
But the other side of the equation is real too: if we cannot find efficiencies and keep the business healthy, there will be fewer movies—and fewer jobs—to protect. The survival of the business is part of the labor question.
I also recognize my bias. My business depends on stars, and I believe audiences will continue to care about human presence and human creativity. I have yet to see much exceptional work emerge from generative AI. The results remain rare and usually come with caveats.
The true believers will push back. If the work is good, audiences will care less and less how it was made. Today’s generative-AI content is the worst we will ever see. The tools will make possible things we cannot yet imagine.
All of that may be true. I still believe human creativity is the decisive ingredient—and assisted AI depends on far more of it.
Who is building assisted-AI tools? What can they actually do? And what can the rest of us learn from outside the room?
Affleck vs. Kavanaugh
That is the question I was chasing when I began looking at Ben Affleck’s InterPositive and Ryan Kavanaugh’s Acme AI & FX.
What I admire is that both of them made large, visible bets on AI-assisted filmmaking before the direction of the market was clear. They won’t be the last. But neither has shown us much about how his system actually works.
I wasn’t looking to dismiss either company. I wanted to understand what problem each was trying to solve—and how hard it would be to get there without access to what it had built.
So I reconstructed both systems, or at least as much of them as I could, from the outside. I gathered interviews, presentations and production claims. I read InterPositive’s patents and compared both companies’ descriptions with existing tools for inpainting, relighting, virtual production, markerless capture, previs, digital environments and footage-conditioned generation.
The question was not whether every component was new. It was how the pieces connected, what information moved between them and whether the resulting system could make hundreds of production decisions repeatable.
The investigation produced some answers, although for the sake of brevity I have not gone as deep here as I would have liked. In broad strokes, InterPositive’s patents reveal its training structure, internal models and the costs it was designed to reduce. Acme’s materials describe a recognizable pipeline that can be compared with tools already on the market.
Neither system has been independently tested. InterPositive has never been publicly demonstrated. Acme has not yet proved that a system built for one unusually ambitious movie—seen only through a short, private work-in-progress presentation—can transfer economically to the next.
Affleck and Kavanaugh were never building rival versions of the same system. They started with different ideas about what should remain fixed, where a director needs freedom and when the machine should enter the process.
Affleck: Build the Model Around the Movie
None of the individual things Affleck says InterPositive can do is entirely new. Visual-effects artists already remove wires, reframe shots, alter lighting, enhance backgrounds and create coverage that was never photographed.
If InterPositive were simply a conversational interface placed over familiar tools, it might be useful. But it would not explain why Netflix paid approximately $587 million for a company with fewer than twenty employees.
So what makes it interesting?
A Film School for the Machine
InterPositive’s most valuable asset may be a library of training footage designed to teach a machine how cinematography works.
Affleck’s team built that library on a soundstage using professional cameras, lenses and LiDAR, which uses lasers to measure the position and distance of objects. They photographed the same scenes repeatedly, changing one variable at a time: the lens, aperture, shutter, sensor, focus or color temperature.
Think of it as a film school for the machine.
Keep the actor, lighting and camera position exactly the same but change the lens, and the machine can begin to understand what that lens does to the image. Change only the aperture, and it can learn how depth of field works.
Footage scraped from the internet cannot teach this nearly as cleanly. An 85mm lens often appears in a portrait, while a wide lens often appears in a landscape or action scene. A model trained on random footage may associate the lens with the subject without understanding how the lens itself changes the picture.
InterPositive isolated the variables so the machine could learn cause and effect.
Training Smart, Not Broad
The footage from a single movie could never contain enough information to teach a model everything it needs to know. It may show one actor in one room under one set of conditions, but not that room from every angle or what would happen if the camera, lens and lighting changed.
Large general-purpose models solve that problem by training on enormous quantities of material. That gives them breadth, but it also creates many of the copyright, consent and chain-of-title problems surrounding AI.
InterPositive appears to take a different approach: train smart rather than broad. Its controlled dataset teaches the general rules of lenses, light, cameras and physical space. The dailies then teach the system the specific world of one movie—its actors, sets and visual language.
One dataset teaches the grammar of filmmaking; the other teaches the vocabulary of this particular film.
The patents leave room for InterPositive to use larger models made by other companies. But if its controlled dataset works as intended, the system may rely less on broadly scraped material because it has been taught the information it actually needs.
Jack of All Trades, Master of… Many?
The patents identify two internal models. The first is called SamildAnach, an old name associated with the Irish god Lugh that roughly means skilled in many arts.
SamildAnach looks at ordinary footage and describes how it was photographed. It tries to identify the lens, camera position, depth, distance and composition. Instead of simply labeling an image “a woman in a room,” it describes how the filmmaker created the shot.
A second model, called Filmmaker, uses those descriptions to modify or generate images while following the visual rules learned from the soundstage footage.
InterPositive is not asking a filmmaker to type “give me a car chase in Tokyo” and accept whatever appears. It begins with a movie’s own footage, learns that production’s performers, sets and visual language, and creates new material within those boundaries.
If it produces an insert or camera angle that was never photographed, it is still generating an image. It is simply generating from the movie rather than from a blank prompt.
How a Production Might Use It
The likely workflow is that a production shoots normally and supplies its dailies. InterPositive analyzes that footage and adapts a larger model to the performers, sets, lighting and visual language of that particular movie.
A filmmaker could then reshape material after the shoot—fixing mistakes, testing new creative choices or making changes that once required a reshoot or a much larger visual-effects effort.
The results would still need review, compositing, color and approval. Nothing I found suggests that filmmakers or visual-effects artists disappear. The tool expands what they can change and when they can change it.
InterPositive therefore appears to be more than a wrapper around existing software, but less than an entirely new form of generative AI. Its value lies in teaching existing technology to understand filmmaking and making it usable inside a professional production.
Why Netflix Dropped Big Bills on This
Netflix may have wanted a system built specifically for filmmakers, trained on material with a clean chain of title and capable of fitting into its existing production infrastructure. A general-purpose model might be more powerful. A model that understands how filmmakers work—and whose training material Netflix can account for—may be more useful.
Not to be cynical, but Netflix is a technology company, and technology companies now appear contractually obligated to mention AI in every earnings presentation. After a quarter that failed to excite investors despite strong earnings, InterPositive gave Netflix something new to talk about.
But $587 million is an expensive way to zhuzh up an earnings call.
Kavanaugh: Build the Movie Before You Shoot It
Kavanaugh begins with the same problem as Affleck—movies cost too much and could benefit from new tools—but his solution is almost the reverse. Acme tries to design the movie so completely in advance that the physical shoot becomes the shortest part of the process.
Its test case is Bitcoin: Killing Satoshi, directed by Doug Liman and starring Casey Affleck and Gal Gadot.
Acme says a conventional version would have required seventy-two shooting days across 213 locations in eighteen countries. Instead, the actors were photographed over approximately twenty days on a gray stage in West London, and the locations were built around them later.
Do the Work Before the Actors Arrive
Kavanaugh says he spent approximately $15 million, seven to eight months and the labor of a fifty-person creative team developing the system and planning the movie.
Every shot, camera position and environment was mapped before production began. The company says its software can read a screenplay and help generate storyboards, blocking plans and three-dimensional previews. Another part of the system determines which props must be physically present because the actors need to touch them. A central planning tool connects the shots with cast schedules, crew requirements and production logistics.
The simple version is that Acme turns the screenplay into a detailed set of instructions before anyone starts shooting. It does not eliminate production decisions. It forces filmmakers to make them earlier.
Why a Gray Stage?
Acme shot the actors in a gray room rather than against a traditional green screen.
Green screens make it easier to separate an actor from the background, but they also cast green light onto skin, hair and clothing. Visual-effects artists then have to remove that color, sometimes frame by frame. A gray stage creates less color contamination and provides cleaner footage for relighting and compositing.
The actors still wear real costumes and handle real props. Multiple cameras record their faces and bodies without motion-capture suits or tracking dots. The director and cinematographer can view versions of the digital environments through augmented-reality equipment, while projected marks on the floor show the actors where walls, furniture and other objects will eventually appear.
Acme is not generating the actors. It is capturing their performances and separating those performances from the worlds around them.
Build the World Afterward
Once the actors are photographed, the production moves into what Acme calls its “VFX Soup.” Production designers, cinematographers, costume designers and visual-effects artists continue building the movie together.
The production designer creates the environments. The cinematographer helps match the light on the actors to the light in those environments. Costumes remain physical but must be adjusted to reflect the colors and light of places that never surrounded them on set.
The advantage is that a director can change those places after the actors have gone home. Liman reportedly turned an ordinary restaurant scene into an elaborate venue with circus performers. A phone call originally set at a desk in Antigua moved to an Antarctic snowstorm with wolves in the background.
The company screened a polished ten-minute compilation shortly after the shoot ended. That does not mean the film was finished in two weeks. It means enough work had been completed during preparation for selected scenes to move through postproduction quickly.
What Does It Save?
Acme reports a total budget of approximately $70 million and a below-the-line cost of $40 million, compared with an earlier physical-production estimate of $120 million to $130 million.
I would want to see the comparison budget. No experienced producer would literally travel to 213 locations. They would be combined, substituted or rewritten. The fair comparison is between Acme’s method and the practical version a conventional line producer would actually budget. Maybe a better comparison is what they get for $25 million (which I believe was the BTL minus the 15 million in product development).
Even with that caveat, shooting for twenty-one days in one building obviously saves money on travel, construction, location fees, company moves and physical-production time. That is Kavanaugh’s point as he sells the system.
So What Is Proprietary?
Acme claims a large patent portfolio—more than sixty patents, compared with Affleck’s four—but I could not verify it. The applications may not yet be public because publication can take eighteen months. That does not mean the patents do not exist.
Most of the individual tools Acme describes already exist. Filmmakers can license software for storyboarding, virtual environments, camera tracking, digital relighting and markerless performance capture. Gray stages and augmented-reality previews are not new inventions on their own.
What Acme may own is the system that connects them.
Think of it as an assembly line. Acme may not have invented every machine on the factory floor—although perhaps it did invent some—but it may have developed a better way to move a movie from script to planning, from planning to the gray stage, and from the gray stage into post without losing information along the way.
What Does a Director Give Up?
The price of Acme’s postproduction flexibility is that more decisions must be made before the shoot.
Actors sitting in a planned wide shot can still improvise dialogue. The problem comes when someone unexpectedly runs across the room and the director whips the camera around to follow. The digital environment on that side of the room may not exist because nobody planned to photograph it.
The system can accommodate changes, but rebuilding the environment, tracking and lighting costs money. At some point, preserving unlimited spontaneity begins to consume the savings the system was designed to create.
That makes the method better suited to some genres than others. A controlled drama or international thriller may benefit enormously. A comedy built around physical improvisation may find it harder to use.
And the biggest question: does it use LLMs? Its hard to tell. But that might complicate chain of title. Apparently some of the cost of producing this movie were the lawyers they needed to avoid that very issue. Who knows…
The Cannes Reaction
What buyers actually saw at Cannes remains fuzzy. Reports describe only “first footage,” elsewhere characterized as a ten-minute preview, shown behind closed doors while the film was still deep in postproduction.
Doug Liman called the response “strong,” and trade coverage noted a “steady stream of curious buyers.” Word on the street was reportedly more mixed, with some attendees said to be underwhelmed.
Because no buyers spoke on the record, it is difficult to know what they saw or how representative those reactions were.
The footage apparently generated some territorial sales. Warner Bros.’ Clockwork reportedly entered exclusive talks for North America, although no completed deal was publicly confirmed.
Can Acme Repeat It?
Kavanaugh’s answer is his slate.
Acme says it has more than fifteen film and television projects in development or production, plus advertising work. None of this has been corroborated beyond Kavanaugh’s claims.
He says the buzz, set visits and conversations around Hollywood have brought producers to Acme with movies they cannot fit into budgets that will get greenlit. The pipeline will now include both Acme’s own projects and production services for others.
It is easy to be cynical about a business whose information comes primarily from its founder. But Kavanaugh took the first steps. He developed a script, got it funded and used that money to build a pipeline that could—or ostensibly does—bring costs down.
He is one giant step ahead of most of us.
Two Different Places for Creative Freedom
InterPositive treats the photographed image—and everything within it—as the source of truth, using AI to alter or extend what was shot. Acme treats the performance as the source of truth, building the surrounding world afterward. One targets reshoots, visual effects and finishing; the other targets locations, travel, construction and shooting days.
The eventual answer probably will not be one system or the other. A filmmaker might shoot 40 percent of a movie through Acme’s process, photograph the rest conventionally and run everything through InterPositive.
That is how movies have always been built: combine tools, move money around and preserve what matters creatively. The technology will change. The job remains deciding where the savings are worth the trade.
The Other End of the Market
Let’s be honest.
Most of us do not operate in Affleck’s or Kavanaugh’s orbit. Wherever the money came from, Kavanaugh put together a $70 million film about a man trying to save Bitcoin. A-list talent or not, he can access capital in a way most filmmakers cannot. And even if you make content for Netflix, who knows when—or whether—you will get access to Affleck’s tools.
Hernany Perla offers a view from the other end of the market. He is an old friend. We met when he was an executive at Lionsgate and the company I worked for had a deal there. After leaving the studio, he became expert at something few people know how to do anymore: take a movie from idea to delivery for a few hundred thousand dollars.
Hart’s company recently brought Perla in to help with microdramas but quickly realized his real skill was stretching one dollar into ten, whatever the format. So it asked him to help supervise Exotic Foods, a show being made for Hart’s LOL platform.
The Work Moves
Exotic Foods was the first project Perla observed that ran entirely through a hybrid pipeline: one location standing in for many, with real actors and props inside a world built around them digitally. Its budget was a small fraction of what full live action would have cost. It was a true indie assisted-AI show.
What he learned was that the savings were not where people assumed. They were not in post. They were in prep, paid for with a level of discipline low-budget filmmaking has spent decades learning to live without.
“You have to pre-plan absolutely everything. Cameras have to match what is going to eventually be replaced by a digital camera. But for it to be seamless, that means you have to lock it, get the lighting—it gets really specific. You have to get the frame rate perfect. The button on your shirt needs to be pre-planned.”
“Dozens of artists doing it, except they get to do it from one location using computers. That’s where the cost savings come in.”
Perla still thinks physically. A table, cart or doorway gives an actor something real to move around. Hair, makeup and costume provide details the eye trusts. The background may be virtual, but the image still needs weight, texture and overlap.
The specific tools matter less than the system around them.
Unreal Engine, already common in production, can build the digital environments. Platforms such as Voia and Beeble help place photographed actors inside those environments by separating, tracking and relighting the live-action footage. Together, they offer pieces of a virtual-production pipeline that once required far more specialized infrastructure.
In the indie world, the producer does not need to own every tool. The job is understanding how they fit together—and where they are likely to fail (and will they work for your budget).
Ambition Leaks Out a Shot at a Time
A low-budget project rarely loses its ambition through one enormous compromise. It happens through dozens of small ones.
For now, that may be where these tools matter most: not replacing the shoot, but intervening in small, controlled ways and solving problems that would otherwise remain in the movie.
Those interventions will grow as the technology improves. Filmmakers will also get better at recognizing where the tools can help. The real skill may not be generating entire worlds. It may be spotting the handful of places where a modest intervention preserves something the budget was about to take away.
Eventually the only restraint will be the filmmaker’s imagination, not the budget.
We are not there yet. But we can start learning where the tools bend, where they break and which compromises they might allow us to stop making.
Who Owns the System?
Perla reached many of the same conclusions as Kavanaugh, only from the other end of the budget scale.
“He’s doing it at the highest level that’s been done yet. I’m looking at probably one of the smaller levels. But what’s ironic is that the principles are the same.”
The difference is ownership.
Affleck built an asset Netflix could own. Kavanaugh built a system Acme hopes to repeat. Perla assembles what a particular production needs, then takes it apart when the movie is finished.
Those are different economic models, but they are also different answers to a larger question: How much of filmmaking should become permanent infrastructure?
Maybe the advantage will belong to the companies that own the technology. Maybe it will belong to producers who know how to combine tools without owning them. Or maybe filmmakers will eventually need some version of both—access to the open market and enough proprietary knowledge to avoid becoming customers in someone else’s system.
Still Digging
I do not have a clean answer yet. I am just digging in.
But here is a story to leave you with.
We were finishing a movie when an AI company approached us. We were missing a few inserts, and several sequences could have used shots we did not have. We reviewed them together, and the company was confident it could help. It promised 4K delivery and indemnification against legal claims, cited deals underway with studios, and pointed to other filmmakers already using its tools.
I was skeptical, but curious.
So I negotiated a very affordable deal. I told our post supervisor and my producing partner that we were paying to learn—but that we should keep pursuing the conventional solution, just in case.
Good thing we did.
The latest update is that we cannot color-correct the company’s outputs. If we want an insert of an actor’s hand holding a phone, we first have to color correct and output the footage in 4k, train the system and then generate the insert. The color has to be baked in before the new shot exists. Several other things we needed—including crowd work—turned out to be beyond what the tools could deliver.
This is a well-known company working in assisted AI. Its limitations do not prove that the technology is useless. They reveal the distance between an impressive capability and something that can survive a real postproduction pipeline.
So how far along are we, really?
Affleck has built technology almost nobody outside a small circle has seen. Kavanaugh has a movie in post that buyers encountered only through a closed, work-in-progress presentation.
Exotic Foods, the show Perla supervised, will soon go directly to consumers. It does not have to survive the same scrutiny as a studio feature projected on a forty-foot screen. Most people will watch it on a phone. It can still generate CPMs, attract brand deals and find an audience. Will those viewers care how the backgrounds were created?
These projects are different and deserve different scrutiny. That may be the point. “Ready” is not a single threshold. A tool can be ready for one format, budget and screen while remaining unusable for another.
Adoption will not happen evenly either. Different countries, platforms and production cultures will balance labor, regulation, cost and quality in different ways. Some may move faster precisely because they are willing to tolerate results that Hollywood cannot—or will not—accept.
The messier tools may reach audiences before the precise ones do.
I am talking to artists and technical experts. I am weighing whether my company should build something of its own—not because I know what it should be, but because ownership may be the price of a seat at the table.
Tomorrow, the tools will change again.





Much more color, answers, definition, examples which stem from the scope of yesterday’s convo. Great piece Ben!