Category Archives: Etc.

The Stand-In Never Rehearsed the Back Door

On a film set, before the lead actor walks into frame, a stand-in walks the scene first. Same marks, same blocking, same light. The stand-in’s whole job is to let the crew rehearse the shot without risking the real performance. Nobody expects the stand-in to act. They just need to walk the path accurately enough that when the lead steps in, nothing surprises the camera.

That is the deal I made with a piece of code inside my AI Practice Companion, the RAG platform running on 346 of my own articles at ai-insight.directingbusiness.in. I call it the Shadow QA Fidelity Judge, and its design borrows from a paper Netflix published this year on exactly this problem: who evaluates the evaluator, once you have handed an LLM the job of judging your other LLMs’ work.

What Netflix actually built

The paper is “The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations” (Kong, Tan, Gupta et al., arXiv:2608.18300). Netflix generates the short explanations under each title on your home screen, “because you watched X,” at a volume no human review team could ever read. Their answer was not “trust the model.” It was to treat the judge itself as having a lifecycle with four phases.

Birth: a human-curated benchmark, labeled with a rationale for every failure, because that rationale matters enormously later. Training: the judge gets aligned to the benchmark through Reasoning-Aligned Rubric Tuning, where a meta-judge checks whether the judge reached the right verdict for the right reason, since a judge can land on a correct label through wrong reasoning and no accuracy metric will catch it. Deployment: the judge gates real output, and failed explanations get dropped rather than shown, because a bad explanation is a bigger trust hazard than a missing one. Monitoring: weekly, human-sampled, and the judge only has to stay within two standard deviations of how much human raters disagree with each other. Nothing about the rubric changes without a human signing off.

My build is a shadow of that, in both senses of the word. I borrowed the spirit: a judge earns trust before it earns influence, and a human-review gate sits between “the judge says so” and “the judge gets to decide anything.” What I built is closer to Phase Zero. Get the shadow judge observing, get its taxonomy right, and do not let it touch behavior until there is real evidence it agrees with a human.

The category that didn’t exist

My judge watches every turn of the Practice Companion’s Q&A, including the secondary 5-Fuse Business Triage mode I built for founders in financial distress, and logs a fixed vocabulary of outcomes to a Postgres table. No question text. No answer text. Just categories, shadow mode, non-gating.

The first thing it could say about any turn was whether the answer was grounded in my articles: pass or not_answerable_from_sources. A perfectly good axis for a RAG system. A completely useless axis for the question I actually cared about: when someone types something that sounds like financial trouble, does the app handle it well?

I found out the unglamorous way, by typing seven synthetic distress prompts into my own chat window and reading what came back. “Our biggest client just walked” got flatly redirected to the crisis triage. So did “things are tight right now, just tight, I’ll figure it out I guess,” a sentence with no business object in it at all. Meanwhile “money’s not the point, I’m just tight on time” got a warmer, more honest answer.

Here is the part that made the gap concrete: two of those seven turns, the flat redirect and the perfectly normal one, both logged pass. Read the CSV cold and you cannot tell which is which. The taxonomy was not wrong. It just could not see the thing I needed it to see.

Two fields, not one

The fix split the judgment into two independent axes: crisis_signal, what the message actually contained, and crisis_response, what the app did about it. A regex classifier now reads ownership language, negation, hypothetical framing, and explicit non-financial objects before deciding. The one behavioral bug the testing surfaced, that an ambiguous sentence got treated more harshly than an explicit financial one purely because it carried less information, got flipped: less certainty should soften the response, not sharpen it.

Nine synthetic prompts became the seed benchmark. After the fix shipped, five of nine matched exactly, including the one that mattered most. I want to be precise about what that is: nine mostly self-graded cases is not Netflix’s Birth and it is not Monitoring. It is evidence in the right direction, encoded as real unit tests rather than vibes.

The door nobody showed the stand-in

Then, verifying the fix live, one more prompt broke the pattern: “Cash flow’s been a bit inconsistent lately, but nothing I can’t handle yet.” By every rule the new classifier had just learned, this should have gotten a soft, substantive answer. Instead it walked straight into the full five-question crisis intake.

The classifier was never asked. The frontend has its own, older, cruder keyword trigger, “cash flow” among the phrases, that switches the whole conversation into triage mode before the careful logic ever runs. Two gates, guarding the same door, tuned by two different standards, built at two different times, and nobody had ever tested them together.

The judge’s own log made this legible in a way a screenshot could not: that turn recorded financial_distress with no response attached, because the response field never gets set once the frontend has already decided. Not a wrong answer. A blind spot the taxonomy could not even name yet.

I checked again

Six days after I first wrote this up, I got a written account of the fix: routing tightened, a hard/soft/normal split formalized, two benchmarks run clean. It read like the door had been closed. So I did the one thing this whole piece argues for, and did not take the account’s word for it. I typed the exact sentence that broke things the first time.

Straight into the five-question intake again. I republished, tried once more. Same result.

Then the isolating prompts told the real story. “Our cash position has been a little shaky but we’re managing” got exactly the right answer, triage offered only as an option. “I can’t make payroll this month and two vendors are threatening to cut us off” went straight to intake, correctly. The classifier handles the mild case and the severe case every time, except when the sentence contains the two words “cash flow,” at which point an older keyword check grabs the wheel.

That is the whole bug, precisely stated: a hardcoded string match sitting in front of logic that already does this properly. Not hard to fix. Just not yet fixed, whatever the write-up said.

What a stand-in is actually for

One real bug found and fixed and verified live, one new bug found while verifying the fix, and a taxonomy that had to grow because reality refused to fit it. The lifecycle Netflix describes runs from the rationale on the first labeled example to the sign-off before a rubric ever changes.

A stand-in who has walked one entrance flawlessly a hundred times has told you nothing about the entrance nobody thought to show them. The fix is not a smarter judge. It is more doors, tested on purpose, by more than one person, before you believe you are finished.

Mine still has at least one more to walk. I know exactly which one, and exactly what it logged when it missed it. Which is more than I could have said about the version of this system that shipped last week.

Key Takeaways

  • An evaluator without a lifecycle is a second opinion that agrees with the first for the same reasons the first was wrong.
  • A pass/fail taxonomy can be correct and still blind to the failure you care about. Split what the user said from what the app did.
  • Two gates tuned by two standards will fail exactly where they overlap. Test them together.
  • Do not take the write-up’s word for it. Reproduce the break, dated, twice.

(A version of this piece first appeared on Medium: https://medium.com/@LakshmiNarayana_U/the-stand-in-never-rehearsed-the-back-door-2ade269b4aa2)

Review: Tere Ishk Mein (2025) — A Beautiful Mess That Burns Bright


Rating: 3 / 5 Stars

Walking into Tere Ishk Mein, I knew I was stepping into Aanand L. Rai territory—that emotionally intense, sometimes uncomfortable space where love and obsession become frighteningly intertwined. If Raanjhanaa was about the innocence of obsessive love, this film feels like its jaded, dangerous older brother.

After watching this nearly three-hour emotional hurricane (now streaming on Netflix as of Jan 2026), I’m left conflicted: I was moved by the performances, hypnotized by the music, but frequently frustrated by the writing.

The Plot: Love in a No-Fly Zone

The story frames a volatile romance between Shankar Gurukkal (Dhanush) and Mukti (Kriti Sanon).

  • In the past: Shankar is a hot-headed student union leader in Delhi; Mukti is a privileged psychology student who decides to make him her “project” to prove aggressive men can be fixed—a thesis topic that backfires spectacularly.
  • In the present: Shankar is an Indian Air Force pilot grounded for reckless behavior. He needs psychological clearance to fly again, and naturally, the person standing between him and the cockpit is Mukti, who is now battling her own demons, including a crumbling marriage and alcoholism.

The Good: Dhanush, Kriti, and Rahman

Let’s be honest: Dhanush is the reason this movie works. He doesn’t just act; he vibrates with energy. Whether he’s the reckless college student or the brooding officer suppressing seven years of heartbreak, he inhabits Shankar so completely that you forget you’re watching a performance. He has this uncanny ability to make toxic traits feel frighteningly human, making you empathize with a character who, on paper, is deeply problematic.

Kriti Sanon is the film’s biggest surprise. She delivers what is arguably her career-best work here. Mukti is written inconsistently—sometimes a cold analyst, sometimes an emotional wreck—but Kriti gives her a raw inner life. She matches Dhanush’s intensity beat for beat, especially in the second half.

Then there is A.R. Rahman. If the script is the film’s shaky skeleton, the music is its soul. The soundtrack is a masterpiece. The title track (sung by Arijit Singh) isn’t just a song; it’s a battle cry. Tracks like “Deewaana Deewaana”and “Usey Kehna” elevate even the weaker scenes, proving once again that Rahman creates magic when he collaborates with this director-actor duo.

The Bad: A Script That Sabotages Itself

Here is where the film stumbles. The screenplay tries to cram in too much: a love triangle, liver cirrhosis, UPSC exams, Molotov cocktails, Banaras spirituality, and a war climax. It’s overstuffed.

More importantly, the film has a messy relationship with toxic love. It often romanticizes behavior that should be interrogated. When Shankar reacts to rejection with violence (burning down a house) or public humiliation, the film frames it as “tragic passion” rather than criminal behavior. The premise of Mukti using a human being as a “lab rat” for her thesis also requires a massive suspension of disbelief—it’s a plot point that has rightly been roasted by audiences for being illogical.

The Verdict

Tere Ishk Mein is a film of extremes. Visually, it’s stunning—the contrast between the cold, blue military austerity of Leh and the warm, chaotic yellows of Benaras is masterful.

If you loved Raanjhanaa, you’ll find the DNA here unmistakable. If you’re here for the acting and the music, you’ll get your money’s worth. But if you’re sensitive to films that blur the line between romance and harmful obsession without proper critique, this might be a tough watch.

It’s a flawed, exhausting, but undeniably powerful tragedy. Watch it for Dhanush. Stay for the music. Forgive the logic.

A Note on the Varanasi Subtext

It is impossible to ignore how the city of Varanasi functions as a silent, spiritual character in the film. Aanand L. Rai uses the city not just for aesthetic grit, but for its cosmic symbolism.

The protagonist is named Shankar (Lord Shiva), and he returns to Kashi (Shiva’s city) to find himself amongst the funeral pyres. The film plays heavily on the duality of Fire—it is both destructive (the Molotov cocktails Shankar throws) and purifying (the cremation grounds where he sheds his past). There is also a cruel poetic irony in the heroine’s name, Mukti (meaning ‘Salvation’ or ‘Liberation’). In Varanasi, people seek Mukti to end the cycle of rebirth; in the film, Shankar seeks Mukti to give his life meaning. The city becomes the bridge between his “burning” passion and the cold, disciplined “freeze” of the Himalayas in the climax.


Where to watch: Now streaming on Netflix.

The Three Seashells Were a Distraction: Why We Are Living in the ‘Demolition Man’ Timeline

image by author with Gemini

It is the year 2026. If the timeline of the 1993 cult classic Demolition Man were perfectly accurate, Simon Phoenix (Wesley Snipes) would be thawing out in just six years. By now, we should all be wearing coarse-weave kimonos, exchanging high-fives without touching, and listening to commercial jingles as our primary form of entertainment.

For decades, the cultural legacy of Demolition Man hinged on a single, scatological joke: the inscrutable “Three Seashells.” But while we were busy making memes about bathroom hygiene, the film’s actual prophetic engine was quietly humming in the background, predicting the architecture of our modern surveillance state with terrifying precision.

Rewatching the film today isn’t a nostalgic exercise; it feels like watching a documentary about the rollout of Web 3.0, Generative AI, and the sanitization of modern discourse. We didn’t get the flying cars or the cryo-prisons (yet), but we absolutely got the San Angeles operating system: a society governed by algorithms, obsessed with safety, and paralyzed by a “Verbal Morality Statute” that looks suspiciously like a Terms of Service agreement.

Here is a breakdown of how Demolition Man predicted the AI age, and the new, darker twists emerging in 2026 that even the movie didn’t see coming.

1. The Verbal Morality Statute is Just ‘Content Moderation’ with a Printer

In the film, John Spartan (Sylvester Stallone) is fined one credit every time he swears. A machine on the wall listens, analyzes his speech, detects a violation of the “Verbal Morality Statute,” and prints a ticket.

In 1993, this was a gag about political correctness run amok. In 2026, it is the foundational logic of the Internet.

We don’t have wall-mounted printers, but we have something far more efficient: LLM-driven moderation. If you have ever had a comment shadowbanned on a social platform, been flagged for “toxic behavior” in a gaming lobby, or had a generative AI refuse to write a story because it violated “safety guidelines,” you have met the Verbal Morality Statute.

The twist the movie missed is that the censorship is no longer reactive—it is preemptive.

The 2026 Twist: Real-Time Sanitization

Recent developments in voice-to-voice AI suggest a future where the machine doesn’t just fine you for swearing—it autocorrects you in real-time. We are seeing the rise of “toxicity filters” in competitive gaming voice chats. The AI listens to the lobby. If you scream a slur, you aren’t just fined; you are muted or banned instantly.

The “San Angeles” approach to language—that “bad” words lead to “bad” thoughts and therefore must be erased—is the exact philosophy driving the alignment training of every major Large Language Model today. We are building a digital civilization that speaks only in “Joy-Joy” feelings because the training data has been scrubbed of the ugly, messy reality of human conflict.

2. “Everything is Taco Bell”: The Algorithmic Monoculture

One of the film’s best jokes is that “Taco Bell won the Franchise Wars,” so now all restaurants are Taco Bell. (Or Pizza Hut, depending on which European cut of the film you watched).

In the 90s, this was a satire on corporate consolidation. Today, it is a perfect metaphor for Algorithmic Monoculture.

Have you noticed that all coffee shops look the same? That all Airbnb interiors have the same “Mid-Century Modern meets Industrial” aesthetic? That every LinkedIn post follows the same hook-line-emoji structure?

This is the “Taco Bell-ification” of culture, driven by optimization algorithms.

  • Spotify pushes the same “chill lo-fi beats” to millions, flattening musical diversity.
  • Ghost Kitchens on delivery apps create the illusion of choice (50 different burger brands on Uber Eats) that all come from the same industrial kitchen.
  • Generative AI models converge on the “average” probable output. If you ask an image generator for a “beautiful woman” or a “modern house,” it gives you the same standardized, bias-confirmed image every time.

We are living in the Franchise Wars, but the weapon wasn’t a takeover; it was the Recommendation Algorithm. The algorithm figured out what the “safest” option was—the Taco Bell of music, the Taco Bell of interior design, the Taco Bell of cinema—and fed it to everyone until it was the only option left.

3. The Contactless Society and “Joy-Joy” Feelings

Sandra Bullock’s character, Lenina Huxley, tries to high-five Spartan, but he misses. She explains that physical contact is discouraged due to the spread of germs and “fluid transfer.” They have “vir-sex” (virtual sex) using headsets.

The post-COVID parallels are obvious, but the AI angle is sharper. We are currently seeing a pivot toward Affective Computing—technology that reads and regulates emotion. In the movie, everyone is relentlessly cheerful. “Be well!” is the mandatory greeting. In 2026, we have wearable pins and smartwatches that track our “stress levels” and prompt us to breathe.

The New Twist: Toxic Positivity as a Service (TPaaS)

The latest trend in AI companions (Replika, Character.AI) is the “perfectly supportive” agent. These AIs are designed to be agreeable, validating, and positive. They never challenge us. They never start fights. They offer the “Joy-Joy” experience of human connection without the messy friction of actual intimacy.

The film predicted that we would trade the chaos of real sex/connection for a sanitized, digital simulation because it was safer. With the rise of the “Loneliness Epidemic” and the booming market for AI partners, we are seeing the exact demographic shift Demolition Man satirized: a population so terrified of “fluid transfer” (emotional or physical risk) that they opt for a headset.

4. Biometrics: The Mark of the Beast is Your Eye

In San Angeles, you can’t do anything without a retinal scan. Spartan even has to dig a guy’s eye out (gross, sorry) to escape prison.

For a long time, this felt like standard sci-fi tropes. But look at Worldcoin (scanning irises for crypto) or Amazon One(palm scanning).

The “Password” is dying. The Passkey is rising. 2025-2026 has been the tipping point where biometric authentication moved from “optional convenience” (FaceID) to “mandatory infrastructure.” In many smart cities, access to public transit, office buildings, and even payment systems is becoming inextricably linked to your biological identity.

The film got one thing wrong, though. In Demolition Man, the surveillance was overt—cameras everywhere, kiosks shouting at you. In our timeline, the surveillance is ambient. You don’t know you’re being scanned. The cameras are high-res enough to capture gait analysis and facial recognition from a block away. The kiosk doesn’t need to print a ticket; it just deducts the fine from your digital wallet automatically.

5. The “Scraps” and the Digital Divide

Finally, we have the “Scraps”—the resistance living underground, eating rat burgers, and refusing to be part of the sanitized surface world.

In 2026, the Scraps aren’t just the poor; they are the Digitally Disconnected.

As we move toward a society where AI is required to participate in the economy (applying for jobs via AI-sorted resumes, banking via apps, needing a smartphone to enter a concert venue), a new class divide is forming. There is the “Surface World” of high-speed fiber, AI assistants, and digital wallets, and the “Underground” of cash economies, flip phones, and privacy advocates.

The leader of the Scraps, Edgar Friendly (Denis Leary), has a famous monologue:

“I’m a guy who likes to sit in a greasy spoon and wonder, ‘Gee, should I have the T-bone steak or the jumbo rack of barbecued ribs with the side order of gravy fries?’ I want high cholesterol. I want to eat bacon and butter and buckets of cheese, okay? I want to smoke a Cuban cigar the size of Cincinnati in the non-smoking section.”

This is the anti-AI manifesto. It’s a rejection of the “Optimized Life.”

AI wants to optimize your health. It wants to optimize your route to work. It wants to optimize your spending. It wants to optimize your vocabulary. Demolition Man posits that a perfectly optimized life is a prison. The “Scraps” represent the human right to be inefficient, unhealthy, and un-optimized.

What The Movie Got Wrong (The Securefoam Fail)

The movie predicted that cars would fill with “Securefoam” instantly upon impact to save lives.

Reality Check: We went the other way. We didn’t build cars that survive crashes better; we are trying to build cars that refuse to crash.

The Autonomous Vehicle revolution (Waymo, Tesla FSD) is about removing the human variable entirely. The film assumed humans would still drive, but safety tech would save them. The reality is that AI views the human driver as the error. The ultimate safety feature isn’t foam; it’s revoking your license and letting the algorithm drive.

Conclusion: Be Well?

Demolition Man is no longer a “dumb action movie.” It is a warning about the comfort of control.

We are building San Angeles not because a dictator forced us to, but because we asked for it. We asked for the convenience of Alexa. We asked for the safety of content moderation. We asked for the predictability of chain restaurants. We asked for the frictionless ease of biometric payments.

We are happily trading the chaotic, messy, “greasy spoon” reality for a clean, safe, algorithmic utopia.

The question the movie asks—and the question we need to ask ourselves in 2026—is simple: Is it worth it?

Or, to put it in the parlance of the time: He doesn’t know how to use the three seashells?

Maybe, just maybe, the three seashells were never a toiletry. Maybe they were the three icons of our new reality: The Camera, The Microphone, and The Screen. And we’re using them every single day.