Category Archives: AI

The Stand-In Never Rehearsed the Back Door

On a film set, before the lead actor walks into frame, a stand-in walks the scene first. Same marks, same blocking, same light. The stand-in’s whole job is to let the crew rehearse the shot without risking the real performance. Nobody expects the stand-in to act. They just need to walk the path accurately enough that when the lead steps in, nothing surprises the camera.

That is the deal I made with a piece of code inside my AI Practice Companion, the RAG platform running on 346 of my own articles at ai-insight.directingbusiness.in. I call it the Shadow QA Fidelity Judge, and its design borrows from a paper Netflix published this year on exactly this problem: who evaluates the evaluator, once you have handed an LLM the job of judging your other LLMs’ work.

What Netflix actually built

The paper is “The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations” (Kong, Tan, Gupta et al., arXiv:2608.18300). Netflix generates the short explanations under each title on your home screen, “because you watched X,” at a volume no human review team could ever read. Their answer was not “trust the model.” It was to treat the judge itself as having a lifecycle with four phases.

Birth: a human-curated benchmark, labeled with a rationale for every failure, because that rationale matters enormously later. Training: the judge gets aligned to the benchmark through Reasoning-Aligned Rubric Tuning, where a meta-judge checks whether the judge reached the right verdict for the right reason, since a judge can land on a correct label through wrong reasoning and no accuracy metric will catch it. Deployment: the judge gates real output, and failed explanations get dropped rather than shown, because a bad explanation is a bigger trust hazard than a missing one. Monitoring: weekly, human-sampled, and the judge only has to stay within two standard deviations of how much human raters disagree with each other. Nothing about the rubric changes without a human signing off.

My build is a shadow of that, in both senses of the word. I borrowed the spirit: a judge earns trust before it earns influence, and a human-review gate sits between “the judge says so” and “the judge gets to decide anything.” What I built is closer to Phase Zero. Get the shadow judge observing, get its taxonomy right, and do not let it touch behavior until there is real evidence it agrees with a human.

The category that didn’t exist

My judge watches every turn of the Practice Companion’s Q&A, including the secondary 5-Fuse Business Triage mode I built for founders in financial distress, and logs a fixed vocabulary of outcomes to a Postgres table. No question text. No answer text. Just categories, shadow mode, non-gating.

The first thing it could say about any turn was whether the answer was grounded in my articles: pass or not_answerable_from_sources. A perfectly good axis for a RAG system. A completely useless axis for the question I actually cared about: when someone types something that sounds like financial trouble, does the app handle it well?

I found out the unglamorous way, by typing seven synthetic distress prompts into my own chat window and reading what came back. “Our biggest client just walked” got flatly redirected to the crisis triage. So did “things are tight right now, just tight, I’ll figure it out I guess,” a sentence with no business object in it at all. Meanwhile “money’s not the point, I’m just tight on time” got a warmer, more honest answer.

Here is the part that made the gap concrete: two of those seven turns, the flat redirect and the perfectly normal one, both logged pass. Read the CSV cold and you cannot tell which is which. The taxonomy was not wrong. It just could not see the thing I needed it to see.

Two fields, not one

The fix split the judgment into two independent axes: crisis_signal, what the message actually contained, and crisis_response, what the app did about it. A regex classifier now reads ownership language, negation, hypothetical framing, and explicit non-financial objects before deciding. The one behavioral bug the testing surfaced, that an ambiguous sentence got treated more harshly than an explicit financial one purely because it carried less information, got flipped: less certainty should soften the response, not sharpen it.

Nine synthetic prompts became the seed benchmark. After the fix shipped, five of nine matched exactly, including the one that mattered most. I want to be precise about what that is: nine mostly self-graded cases is not Netflix’s Birth and it is not Monitoring. It is evidence in the right direction, encoded as real unit tests rather than vibes.

The door nobody showed the stand-in

Then, verifying the fix live, one more prompt broke the pattern: “Cash flow’s been a bit inconsistent lately, but nothing I can’t handle yet.” By every rule the new classifier had just learned, this should have gotten a soft, substantive answer. Instead it walked straight into the full five-question crisis intake.

The classifier was never asked. The frontend has its own, older, cruder keyword trigger, “cash flow” among the phrases, that switches the whole conversation into triage mode before the careful logic ever runs. Two gates, guarding the same door, tuned by two different standards, built at two different times, and nobody had ever tested them together.

The judge’s own log made this legible in a way a screenshot could not: that turn recorded financial_distress with no response attached, because the response field never gets set once the frontend has already decided. Not a wrong answer. A blind spot the taxonomy could not even name yet.

I checked again

Six days after I first wrote this up, I got a written account of the fix: routing tightened, a hard/soft/normal split formalized, two benchmarks run clean. It read like the door had been closed. So I did the one thing this whole piece argues for, and did not take the account’s word for it. I typed the exact sentence that broke things the first time.

Straight into the five-question intake again. I republished, tried once more. Same result.

Then the isolating prompts told the real story. “Our cash position has been a little shaky but we’re managing” got exactly the right answer, triage offered only as an option. “I can’t make payroll this month and two vendors are threatening to cut us off” went straight to intake, correctly. The classifier handles the mild case and the severe case every time, except when the sentence contains the two words “cash flow,” at which point an older keyword check grabs the wheel.

That is the whole bug, precisely stated: a hardcoded string match sitting in front of logic that already does this properly. Not hard to fix. Just not yet fixed, whatever the write-up said.

What a stand-in is actually for

One real bug found and fixed and verified live, one new bug found while verifying the fix, and a taxonomy that had to grow because reality refused to fit it. The lifecycle Netflix describes runs from the rationale on the first labeled example to the sign-off before a rubric ever changes.

A stand-in who has walked one entrance flawlessly a hundred times has told you nothing about the entrance nobody thought to show them. The fix is not a smarter judge. It is more doors, tested on purpose, by more than one person, before you believe you are finished.

Mine still has at least one more to walk. I know exactly which one, and exactly what it logged when it missed it. Which is more than I could have said about the version of this system that shipped last week.

Key Takeaways

  • An evaluator without a lifecycle is a second opinion that agrees with the first for the same reasons the first was wrong.
  • A pass/fail taxonomy can be correct and still blind to the failure you care about. Split what the user said from what the app did.
  • Two gates tuned by two standards will fail exactly where they overlap. Test them together.
  • Do not take the write-up’s word for it. Reproduce the break, dated, twice.

(A version of this piece first appeared on Medium: https://medium.com/@LakshmiNarayana_U/the-stand-in-never-rehearsed-the-back-door-2ade269b4aa2)

The Devil Still Wears Prada. The Plot, Unfortunately, Shows Its Seams.

Twenty years later, Miranda Priestly is still formidable. I’m less convinced about the story built around her.

I caught The Devil Wears Prada 2 recently on JioHotstar.

Maybe that was the right way to watch it.

No opening-weekend hype. No need to immediately decide whether a sequel arriving two decades later had justified its existence. Just the curiosity of meeting some very familiar characters again.

And, if I am being truthful, one character in particular.

I wanted Miranda Priestly back.

Not a softened Miranda. Not a more understanding Miranda shaped by two decades of changing workplace culture.

I wanted the intimidating, immaculately controlled Miranda who could ruin somebody’s day by barely changing her tone.

On that count, the movie delivers.

Miranda is still Miranda

Meryl Streep slips back into Miranda Priestly almost too easily.

The pauses still work. The looks still work. The voice still works.

Most importantly, Streep remembers what made Miranda such a terrific character in the first place: she never needs to demonstrate power. Everyone around her demonstrates it for her.

And yes, I rather enjoyed seeing badass Miranda again.

One of her sharper lines in the sequel is:

“You’re not a visionary. You’re a vendor.”

That is classic Miranda. No speech. No elaborate insult. Seven words and somebody’s self-esteem has left the building.

And, of course:

“That’s all.”

Some lines simply survive their movies.

The problem is that while Miranda appears completely effortless, the plot around her often feels exactly the opposite.

I could see the writers arranging the pieces

Twenty years have passed.

Andy Sachs is no longer the unsure young assistant entering a world she doesn’t understand. Emily has moved on. Nigel has his own journey. The fashion and publishing industries themselves have changed enormously.

There is actually a very good reason for a sequel to exist here.

Traditional magazines are no longer operating from the position of cultural dominance they enjoyed when the first movie came out. Audiences have fragmented. Advertising has shifted. Influence has moved elsewhere.

That could have been the movie.

Instead, I often felt the writers first decided they wanted Miranda, Andy, Emily and Nigel back in each other’s lives, and then worked backwards to create reasons for it.

Characters need to meet, so circumstances emerge.

Old tensions need to return, so situations are constructed to bring them back.

Andy needs to re-enter Miranda’s universe, so the screenplay finds a route to get her there.

None of it is impossible.

It just feels a bit forced.

And at times, contrived.

That was my biggest issue with the movie.

But nostalgia works

The funny thing is, I didn’t mind all that much while these characters were actually on screen.

There is genuine pleasure in seeing them again.

Anne Hathaway plays Andy as someone who has clearly lived a life since we last saw her. Stanley Tucci remains instantly watchable. And Emily Blunt gets some of the movie’s most enjoyable moments.

Emily probably gets the most quotable line too:

“May the bridges I burn light my way.”

Perfectly Emily.

The sequel also knows exactly how much affection audiences have for these characters.

Perhaps a little too well.

There are callbacks, familiar rhythms and scenes designed to make us think:

Ah, yes. I remember this.

And most of the time, it works.

I smiled.

But occasionally I became aware that I was smiling because the movie was reminding me of something I had already loved, rather than giving me something equally memorable now.

That is the danger with sequels like this.

The first movie had to make us care.

The second one already knows that we do.

The part I wanted more of

The most interesting aspect of The Devil Wears Prada 2 for me was not the relationships or the nostalgia.

It was the changing media business.

The original movie showed Runway as an institution at the height of its power. Miranda was formidable partly because Runway itself mattered enormously.

The sequel is set in a very different world.

And that raises a much more interesting question:

What happens to powerful people when the institutions that created their power start losing relevance?

That is where Miranda becomes fascinating again.

She may still dominate a room, but does the room itself matter as much as it once did?

For someone who has spent years around media and digital businesses, I found myself wanting the film to go much deeper into that tension.

Magazines once decided what mattered.

Now influence flows through creators, platforms, algorithms, communities and fragmented audiences.

So what does somebody like Miranda do when cultural gatekeeping itself is changing?

That, to me, was the sequel hiding inside the sequel.

I would happily have traded some of the manufactured plot complications for more of that.

So, did I like it?

Yes.

But with reservations.

It is stylish, funny and entertaining enough.

The cast still has chemistry. The fashion world remains fun to revisit. And whenever Meryl Streep appears, the movie immediately becomes more interesting.

The problem is that I could sometimes see the machinery underneath the story.

The original Devil Wears Prada felt like a strong story that happened to give us unforgettable characters.

This one occasionally feels like the filmmakers first decided to bring back unforgettable characters and then went looking for a story.

Fortunately, they remain very good company.

And Miranda Priestly?

Still intimidating.

Still funny.

Still able to make everyone around her look slightly nervous without appearing to do anything at all.

I enjoyed seeing her again.

I just wish the plot had relaxed a little and allowed her—and the changing world around her—to do more of the work.

My rating: 3.5/5

Good to see them again.

Very good to see her again.

That’s all.

House of the Dragon Season 3: The Throne That Corrupts Absolutely

I wrote a while back about how Yellowstone is The Godfather on horseback — a story about what an empire does to the people defending it. House of the Dragon Season 3 just did the same trick with dragons, and did it more deliberately than any season of this franchise so far.

Season 3 wrapped on August 9, 2026, eight episodes after it began, and it is the best-reviewed run House of the Dragon has had — a Certified Fresh 95% on Rotten Tomatoes (its highest of the three seasons, matching the two best seasons of the original Game of Thrones), a 76/100 “generally favorable” on Metacritic, and a premiere block that set IMDb episode-rating records for the franchise (9.2 for episode one, 9.4 for episode two). The numbers back up what watching it feels like: this is the show finally being precise about what it’s always been about.

⚠️ SPOILER ALERT: everything below assumes you’ve finished the season, including the finale. Turn back now if you haven’t.

What the season is actually about

Every season of this show gestures at power and its costs. Season 3 is the first one that says it outright. Showrunner Ryan Condal has confirmed the season was built around Lord Acton’s 1887 line — power tends to corrupt, and absolute power corrupts absolutely — and around a single driving question he kept asking the writers’ room: what does the throne actually do to you, the closer you get to it?

That’s not a subtext I’m reading into it. It’s the stated brief. And the season stages the answer methodically, mostly through one character.

Rhaenyra’s fall is the whole season

Rhaenyra opens Season 3 the way she’s opened every season — deliberate, principled, trying to prove a woman can rule without ruling like the men before her. She ends it ordering deaths, spiraling when Aegon II turns up alive with Sunfyre, and asking Alicent — her oldest friend turned oldest enemy — to help kill Alicent’s own son, Aemond. Condal has since compared her trajectory directly to Daenerys Targaryen’s fall in the original series: a leader convinced she’s owed the throne by right, backed by overwhelming force, running out of people willing to tell her no.

The hinge is Helaena’s death partway through the season. Emma D’Arcy has described Helaena as “the only character who has proximity to power but remains kind of uncontaminated” — close enough to see everything, corrupted by none of it. Her death is the moment, per D’Arcy and Condal both, that “condemns Rhaenyra to the path she’s chosen.” There’s no version of the finale where Rhaenyra comes back from that, because there’s no one left in the room whose disappointment still costs her anything.

Then Aegon survives. The killing that was supposed to settle things settles nothing, and that failure — not a villain twist, just the plan not working — is what tips her into the “unhinged tyranny” the finale is built around.

The court fractures right along with her

The detail I liked most, on a rewatch: D’Arcy and Condal have both said Rhaenyra’s three closest relationships — Daemon, Mysaria, Alicent — function as externalized pieces of her own splintering judgment across the season. Daemon is force. Mysaria is cunning. Alicent is restraint. Instead of integrating the three into one decision, she starts acting on whichever voice got to her most recently.

Which sets up the season’s best structural swerve: Daemon — established across two seasons as the impulsive Rogue Prince — becomes, by the midpoint, the calmest, most clear-eyed person on Team Black, exactly as Rhaenyra absorbs the volatility he used to carry. They basically swap temperaments. It’s a smarter piece of writing than “the reckless one finally grows up” — it plays more like stability is a role the system needs filled, and it just relocated to whoever was steadiest that week.

Aemond and Alicent: proximity without the actual seat

Aemond spends most of the season parked at Harrenhal — a deliberate echo of Daemon’s own Season 2 exile there — which is also the season’s most-criticized pacing choice; a lot of reviews wanted more of him and got Alys Rivers instead. He resurfaces alive in the finale, dragon intact, glimpsed on the Iron Throne in a flash-forward clearly meant for Season 4.

Alicent’s arc is messier by design, and critics are right that some of it plays abrupt — her ride to Harrenhal to move against her own son is the most-argued-about beat of the back half. But there’s a reading of it that’s more interesting than “bad plotting”: three seasons adjacent to the throne, never quite holding it, and when the legitimate channels run out, she acts unilaterally and badly. That’s what influence without an actual seat tends to do, eventually.

What worked and what didn’t

What worked:

  • Condal finally states the thesis instead of implying it, and the season is sharper for having one clear question to answer
  • Emma D’Arcy’s performance carries the entire back half — Rhaenyra’s unraveling never once feels like a plot mechanic
  • The Daemon/Rhaenyra role reversal is genuinely clever writing, not just a twist for its own sake
  • Three real set pieces (the Battle of the Gullet, the Storming of the Dragonpit, the Battle of Tumbleton) that each escalate the argument instead of just being spectacle

What didn’t:

  • Aemond gets sidelined at Harrenhal for most of the season, and the show knows it — several reviews single this out as the season’s biggest pacing problem
  • Alicent’s turn against Aemond needed another episode or two of setup; it reads as sudden even when the underlying logic holds up
  • Helaena’s death, however thematically necessary, is dispatched with less weight than a moment that important deserved

My verdict: 8.5/10. The most purposeful season of this show yet — it knows exactly what it’s arguing and proves it scene by scene, even where the plotting gets ahead of itself.

What’s next

Season 4 is confirmed as the final chapter, expected in 2028. Given where this one left Rhaenyra — and given Condal is openly inviting the Daenerys comparison — I’d bet the endgame is less “who wins the Dance of the Dragons” and more “what’s left of her by the time it ends.” Worth staying for.

Have you finished Season 3 yet? Curious whether you’re Team Rhaenyra-had-no-choice or Team she-lost-the-plot-completely — drop it in the comments.


Sources: Deadline, Variety, ScreenRant, Collider, MovieWeb, Forbes, Rotten Tomatoes, Metacritic, The Bull’s Eye.

Categories: 1-By Laksh, ET, Movies, TV