Vishnu Yash Pandey

I build AI products from messy problems to working, tested workflows.

I’m interested in what happens between a model’s output and a user’s decision — where evidence, deterministic logic, human review, and real-world failures shape the product.

The interesting part happens after the model works.

Anyone can pipe a prompt through an API and produce plausible prose. The product work begins when you must decide where AI actually belongs, where deterministic logic is safer, how evidence is surfaced to a skeptical operator, and how to gracefully absorb failure when the model is wrong.

Noto

“A preference looked like a decision.”

Sara preferred 15 minutes, Devraj preferred 5, Marcus said “let's test both before deciding.” Noto initially collapsed conversational exploration into a finalized, authoritative decision.

Ground candidate decisions directly to transcript spans and require explicit human review before committing them.
SquadPay

“AI passed reconciliation and was still wrong.”

Receipt extraction produced numeric line-item values roughly 1,000× too small, but they were consistent with each other, so SquadPay's own reconciliation check passed. Agreeing numbers aren't the same as correct numbers.

Keep transaction arithmetic deterministic and check for anomalies: unusual decimal precision, suspiciously low totals, non-positive totals.
NetSense AI

“A healthy score sat next to an active high-severity incident.”

A point-in-time composite anomaly score read 0.16 (Normal), while an oscillating, unstable corridor experienced an active, high-severity BGP Peer Dampening incident right beside it.

Future direction: rolling-window / variance-aware scoring (not yet implemented).
— Those failures changed the products.

Selected Work

Noto

— AI Meeting → Execution Copilot

Noto turns messy meeting transcripts into evidence-backed, reviewable execution items — with human approval before decisions and actions become authoritative.

The Failure

“A preference looked like a decision.”

Sara preferred a 15-minute slot, Devraj preferred 5 minutes, Marcus said “let's test both before deciding.” Noto initially collapsed conversational exploration into a finalized, authoritative decision.

18 transcripts · 54 ground-truth items · 78.0 F1 (best prompt-only version)

SquadPay

— Making Settling Up Less Awkward

A shared-expense workflow where AI extracts messy receipt data, while deterministic logic handles the money — not a fintech product, an AI-assisted expense product.

The Failure

“AI passed reconciliation and was still wrong.”

Extracted values were roughly 1,000× too small, but they were consistent with each other, so SquadPay's own reconciliation check passed. Agreeing numbers aren't the same as correct numbers.

15 evaluated receipts · 69/69 product tests passing · Gemini 3.6 Flash

NetSense AI

— Network Fault Detection & Incident Investigation

A network-operations prototype connecting telemetry anomalies to the evidence and incident context engineers need to investigate them. Deterministic, rule-based scoring on synthetic telemetry.

The Failure

“A healthy score sat next to an active high-severity incident.”

A point-in-time composite anomaly score read 0.16 (healthy), while an oscillating, unstable corridor experienced an active, high-severity BGP Peer Dampening incident right beside it.

6 synthetic telemetry profiles · 5/5 resolvable-scenario agreement

I build it, break it, and check what changed.

I’m a final-year CS student at Amity University Lucknow, focused on AI product management. I’ve built three working products, and I test each one until something breaks, change the product, then check whether the change helped. Here’s how that has gone.

Where AI belongs

I decide where the model stops. In Noto it proposes and a person approves. In SquadPay it reads the receipt and deterministic code handles the money. One receipt still reconciled while 1,000× too small, so SquadPay now checks for unusual totals.

Evidence before claims

Each number says what it can't prove. Noto's guardrail passed a targeted retest, but I haven't rerun the full benchmark, so I claim no gain from it. SquadPay's second run gained one receipt from a rate-limit recovery, so I don't credit the model.

Real users

Noto's benchmark missed the failure that changed the product. Eight people testing it found it. In one test meeting, two people preferred different options and a third said “test both before deciding.” Noto initially treated that discussion as a decision when nothing had been decided.

Failures stay visible

NetSense scored a corridor healthy during an active incident, and I left it visible. Retuning the threshold would have hidden a blind spot in point-in-time scoring. Rolling-window scoring is future direction, not built.

Resume

Experience & Education

Experience

Tata Teleservices Limited

AI Product Manager Intern

May–Jun 2026 · Noida
  • Competitive benchmarking
  • Data-backed product & client pitches
  • PRDs, user stories & acceptance criteria
  • GenAI exploration
  • Network-operations exposure

Selected AI Product Work

Noto

AI Meeting → Execution Copilot

18 transcripts · 54 ground-truth items · 78.0 F1

SquadPay

AI-Assisted Bill Splitting

15 evaluated receipts · 69/69 product tests

NetSense AI

Network Operations Prototype

6 telemetry profiles · 5/5 resolvable scenarios

Education

Amity University Lucknow

2023–2027

B.Tech — Computer Science & Engineering · CGPA 7.10

City Montessori School, Lucknow

2021

ISC — PCM · 90%

Certifications

IBM AI Product Manager · EA Product Management · Google Agile PM · Datacom Partnering with AI · Adobe University Hackathon 2026

Product & AI

AI Product Management · Product Discovery · Product Strategy · PRDs · AI Evaluation · LLM Evaluation · User Testing · Guardrails · Human-in-the-Loop · Generative AI · Agile

Let’s build something useful.

Building AI products where model output meets real-world decisions.