In January 2026 the Game Developers Conference found that 64% of people working in visual and technical art hold unfavourable views of generative AI, the most negative of any discipline it surveyed. That is from the GDC industry survey , and it is the correct place to start an article about AI character animation.
The scepticism is not ignorance. Animators have watched a decade of tools promise to remove the tedious part of the job and instead remove the interesting part while leaving the tedium intact.
This is an honest account of where generated motion actually helps, where it does not, and why the boundary between those two sits exactly where it does.
Volume, speed and the willingness to try a fourth idea. Those three, and they are more valuable than they sound.
The pattern is consistent across every tool in this category. Gross body movement transfers well. Stride, weight transfer, the overall shape of an action, coarse timing. That is enough to answer questions, which is what most early stage animation work is for.
One creator working with generated motion described the change in iteration count directly, saying they now lock a motion in about three samples where the same result previously took fourteen to twenty passes. That is the shape of the benefit, and notice what it is not. It is not fewer animators. It is more attempts per animator.
Everything that carries intent. The list is stable, well understood, and has not moved much in two years despite the release notes.
There is a measurable reason contact frames land softly. Peer reviewed comparisons in Sensors in June 2026 put markerless capture within roughly two to three degrees of marker based systems on hip and knee flexion, while secondary rotations degraded to six degrees and beyond. The readable motion survives. The specificity does not.
Every shot has a point where generation stops paying and craft starts. Finding it early is the whole skill of working this way, and four questions locate it reliably.
Answer those four and you know how to budget the shot before touching it. Most shots in a project sit comfortably on the generation side, which is exactly why the ones that do not deserve the time you saved.
Take a character vaulting a low wall and landing badly. Here is what survives generation and what gets replaced.
PINOC’s own capture-to-export walkthrough, posted by @Viggle_PINOC on X.
The generated pass gives you the approach run, the plant, the vault arc and the rough recovery. Timing at a coarse level is right. The character reaches the wall when it should and lands roughly where it should. That is perhaps seventy percent of the shot, arriving in a minute.
What needs replacing is specific and predictable. The plant frame where the hand meets the wall, which will be floating by a few centimetres. The landing contact, which will be soft where it should be heavy. And the two frames of the stumble, which is the entire point of the shot and the reason anybody is watching.
Total hand work, a handful of poses instead of a full blocking pass. That is the actual arithmetic of this way of working, and it explains why the benefit shows up as iteration rather than as fewer people.
PINOC is built around producing a first pass you refine, and the design choices follow from that rather than from trying to produce a finished shot.
On why AI does not replace the animator, from DANI (@alittledani). View the original reel on Instagram .
Three of those choices matter for control specifically. Every run returns several takes rather than one, so you are choosing a performance rather than accepting one. You preview on a character in the viewport before exporting, so nothing enters your scene without being judged. And the export is editable animation data rather than video, so every decision downstream remains yours.
Underneath sits JST, Viggle’s in house model, trained with physical priors rather than pure pose data. The practical effect for an animator is less correction work on weight and ground contact, which is the least interesting part of the cleanup.
What it does not do is finish anything. There is no polish pass, no arc refinement and no timing judgement. The text to motion route in particular produces a starting point that assumes somebody competent is about to take over.
No, and the measured evidence is less flattering to the tools than either side of the argument suggests.
McKinsey’s January 2026 analysis of film and television production found leaders reporting productivity increases of five to ten percent in specific use cases. Not a transformation. A useful, incremental gain concentrated in particular tasks.
That matches what practitioners describe. The work that disappears is the twelfth variation of a background shuffle. The work that remains is everything requiring a decision, and that work does not get faster because the base layer improved.
One opinion stated without hedging. The animators who are getting the most from this are not the ones who trust it most, they are the ones who are hardest to satisfy. Scepticism is a productive stance here, because the tool rewards someone who knows precisely what is wrong with a take and can fix it in four poses. It punishes anyone who cannot tell.
Not on current evidence. Generated motion is reliable on gross movement and unreliable on contacts, arcs, exaggeration and anything carrying intent, which is precisely the work animators are employed for. What changes is the entry point, since refining a base layer is different work from keying from zero. Measured productivity gains in film and television sit at five to ten percent in specific use cases rather than across the board.
On a typical shot, the approach, the overall arc and the coarse timing survive, while contact frames and any performance beat get replaced. That works out to keeping most of the motion and rebuilding the moments that matter, which is a handful of poses rather than a full blocking pass. Hero shots keep far less, sometimes nothing beyond reference.
Because single camera capture infers ground contact rather than measuring it, so contact frames are estimated and land softly. Models trained with physical priors handle this better than pure keypoint tracking, but none eliminate it. Test any tool on a jump landing rather than a walk cycle, since that exposes weight handling within a couple of seconds.
Yes, and it suits experienced animators better than beginners, which is the opposite of what most people assume. The skill it rewards is knowing exactly what is wrong with a take and fixing it efficiently, which is precisely what hand animation teaches. Someone who cannot identify the flaw in a generated pass will ship it, and it will look generic.
Yes, through retargeting, and how much work that is depends on skeleton naming. Motion arriving on a widely used convention maps almost automatically. Some tools now let you upload your mesh and retarget inside the tool, which removes the export and reimport loop entirely. Test this before committing, because it is the largest hidden cost in the workflow.
That depends on your contract and your relationship, but the practical answer is that transparency avoids a worse conversation later. Many studios now ask directly. Framing it accurately helps, since a generated first pass refined by an animator is a different claim from a generated final, and the second one is rarely what anybody actually shipped.
The version of this technology that deserves the scepticism is the one that promises finished animation. That version does not exist, and the tools claiming otherwise are demonstrating on walk cycles.
The version worth using is narrower and duller. It gives you a plausible base layer in a minute so you can spend your judgement on the four frames that carry the shot.
The machine can produce movement. Deciding what that movement means is still the job, and it is still yours.