Hiring AI Engineers (NYC) @Pendo: Vintage data. Fresh AI

:rocket: Hiring @Pendo: Staff and Sr AI Engineers - NYC (onsite)

If you’ve built production AI systems at a startup and care as much about evals and observability as the model itself, this one’s for you

We’re building Novus at Pendo, a product agent that reads live codebases, detects real user pain, and suggests fixes autonomously. Not a demo. Not a prototype. A production system catching ~60% of AI issues before customers ever notice.

What you’d work on:

→ Agent architecture with a sandboxed harness reading and writing against live repos

→ Near data-science-grade queries over real-time behavioral data (trillions of signals)

→ Evals-first culture: LangSmith, two dedicated eval sets, shipped and measured daily

→ Genuinely open problems: product memory, unbounded use cases, proof-of-concept to production

How the team runs:

→ ~3-day sprints. Learn in the morning, ship the same day.

→ Novus improves Novus. Tight feedback loops, always.

Backed by Series F. Led by a CAIO who scaled an AI firm to $100M partnering with OpenAI and LangChain.

If this sounds like your kind of problem, or you know someone it might, drop a reply or DM me.

This is exactly the kind of problem I’ve been solving this week. Built an evals-first benchmarking system that caught a blind spot in model-as-judge evaluation — the quality axis humans prioritize is invisible to automated raters. Two live demos you can run right now:

The Honored Ask — prompt phrasing produces larger quality shifts than model scaling. 9 models, 4 companies, $0.02 total compute.

The Blaine Test — tone calibration benchmark. Can your AI match conversational register? None of the 6 models we tested could.

Data, code, methodology on github but I have to post a separate link since I can’t reply to your message with more than two links. Please check below.

Built the harness, ran the evals, shipped the demos, published the dataset — in one night. “Evals-first culture” isn’t a bullet point for me, it’s how the week went.

Additionally, please check my space and submissions to the HuggingFace Small Build Hackathon. 22 apps timestamped and shipped during that time period.

I look forward to hearing from you.

Kory.Indahl@gmail.com

The Blaine Test — tone calibration benchmark. Can your AI match conversational register? None of the 6 models we tested could.

Data, code, methodology: GitHub - claude-wayfinder/honored-ask-blaine-test: The Honored Ask + The Blaine Test: Two evaluation dimensions your benchmarks are missing. 9 models, 4 companies, 1220 data points, $0.02. · GitHub

Built the harness, ran the evals, shipped the demos, published the dataset — in one night. “Evals-first culture” isn’t a bullet point for me, it’s how the week went.

Additionally, please check my space and submissions to the HuggingFace Small Build Hackathon. 22 apps timestamped and shipped during that time period.

I look forward to hearing

Thank you. I can’t send a DM because your profile is private and I’m limited by what I can post in response here.