Creative Duel: Human Judgment versus AI Assistance

The second Creative Duel of Nivan Live 2026 was the one the room had been waiting for. Bennett Hornbostel, principal UX researcher at Amazon, took human judgment. Varun Kapoor, lead product designer at Microsoft, took AI assistance. The moderator noted the matchup out loud partway through. What emerged was not a disagreement about capability but a disagreement about vocabulary, and the vocabulary turned out to matter.

Creative Duel Human Judgement vs. AI Assistance

The question the format was really asking

The framing question was which is more dangerous, AI being wrong or humans trusting it too quickly. Both debaters treated that as a question about consequences rather than accuracy, which is where the interesting argument lives. Nobody disputed that these systems are capable. The dispute was about the gap between a system producing an answer and a person accepting it.

The disagreement was about one word

The sharpest moment of the duel was lexical. Bennett argued that automation is not judgment, that a system replicating decisions at scale is still automation regardless of how sophisticated it becomes. Varun countered that judgment is the wrong word entirely and that decision is the right one, because automation exists precisely to replicate human judgment and the people resenting it are usually the ones making poor decisions about how to deploy it. That exchange is the whole debate compressed, and neither position resolves it.

Where they actually agreed

Under pressure Varun conceded the structural point cleanly, which strengthened rather than weakened his case. Nothing happens without human involvement. A model is created by a human, produces an answer, and a human accepts or dismisses it. Bennett accepted that framing. The remaining disagreement was about how much value sits inside that loop, and whether expanding what one person can do inside it is a change in kind or only in degree.

The case for human judgment, argued by Bennett Hornbostel

Bennett's argument was that meaningful control requires understanding. He read a formulation from the Notre Dame Institute for Ethics and the Common Good, that when humans approve decisions they did not shape and do not fully understand, they do not have meaningful control but only symbolic approval. He then flipped it into a builder's frame, which is the more useful direction. What you create is only as good as the thinking and judgment humans put into it. His supporting observation is that judgment is already layered invisibly through every product, in the people who envisioned it, the people who built it, the user's ability to understand those decisions, and the user's own subject matter expertise.

Why expertise is doing the work

His clearest example was personal. Asking a research tool whether a customer raised a particular issue, he can immediately recognise when the response misrepresents a participant, because he ran that interview. The moment a product manager interrogates the same system with different expertise and different questions, the context that made his judgment reliable is gone. He extended the logic to training data, noting that companies collecting driving data make invisible judgment calls years earlier about whether to capture snow, and a vehicle trained without it will one day encounter snow with no idea what it is seeing.

The case for AI assistance, argued by Varun Kapoor

Varun opened with a computationally designed rocket engine, printed with an efficient cooling system in a fraction of the time a human team would have required, and drew a simple conclusion. The peer is available, so use it. His most provocative claim is that innovation follows human laziness, and his more precise version of it is worth separating out. One kind of laziness is trivial. The other is choosing what to be relieved of. If he can orchestrate an agent that calls other agents, time leaves his plate and is reallocated to ideas rather than repetition, which makes it a resource allocation decision rather than an embarrassing one.

Hallucination as a design failure

Varun did not dispute that hallucination is real. His position is that it occurs where guardrails are absent, and that setting those guardrails is a human responsibility rather than a reason to withhold trust. His framing was that we are the puppet master, these are our agents, and knowing which repetitive tasks to hand off is the skill. Where Bennett argued that context is lost when someone other than the expert interrogates a system, Varun's counter was that context is something you supply. Set it properly and the output improves. That is a stronger claim about human agency than the human judgment side was making, which is the paradox at the centre of this duel.

What both of them expect next

Asked how this conversation sounds in a year, Bennett predicted it sounds much the same, added that he finds parts of the current wave genuinely unimpressive, and expects more discussion of trust and of training data generated intentionally rather than scraped. Varun predicted Bennett would come round, then made a testable claim about this very event, that the effort behind Nivan Live could be cut by fifty to seventy percent with specialised models and automated distribution. Both landed on the same open question, which is what a junior role becomes when a junior can build the thing that does their job.

What the room decided

No winner was declared. The moderator, having spent the session prodding both, admitted to being on team AI and then asked the room whether anyone was honestly not already working with it. The applause split without resolving. That outcome is a fair reading of where practitioners actually sit. Almost everyone is using these systems daily while remaining unwilling to concede that judgment has moved.

Questions this leaves open

If you approved something this week that you did not fully understand, was that control or was it symbolic approval. When an AI output looks right to you, are you evaluating it or recognising it, and can you tell the difference in the moment. If context is something you supply, as Varun argued, who in your organisation is accountable when the context supplied was wrong. And if a junior employee can build the agent that performs the junior role, what exactly is the apprenticeship that produced people like Bennett, and what replaces it.

The claim worth revisiting in 2027

Bennett's closing question was the sharper of the two. All of this exists to serve human needs, so if you are not applying judgment, what was the point of building it. Varun's closing was about surface area, that AI lets one person do the work of three and earn a better seat at the product table. One of those claims is testable next May, when this room reconvenes and can check whether the effort behind the event actually fell by seventy percent.

Previous
Previous

Mr. Parts: What a Wrong Delivery Taught Us About Designing for People Who Cannot Wait

Next
Next

Why I Started Nivan: Looking at Innovation Through the Lens of Time