Phenom
Applied AI, built to give evaluators candidate insights for faster, more informed decisions.
John SmithNDA Protected — Product Manager
Laura SmithNDA Protected — UX Designer
Adrian Burlău — Product Designer
John SmithNDA Protected — Product Manager
Laura SmithNDA Protected — UX Designer
Adrian Burlău — Product Designer
Disclaimer: This project is under NDA, the screens are recreations for portfolio purposes not the shipped product, but reflective of the actual design solutions I developed.
Overwhelmed evaluators, rushed calls
At Phenom, I helped shape an AI-native feature that reads a candidate's screening interview and gives evaluators structured insights to work from, without replacing their own judgment. I owned the working prototype and AI-output testing, using four recruiter personas to evaluate how different AI approaches behaved before we committed to a direction.
Fair evaluations, just not enough time for all of them
Manual scorecards worked at low volume. At hundreds of candidates per role, evaluators didn't have time to give every answer a fair, in-depth read and judgment quality varied day to day.
Of recruiters and hiring managers struggle with candidate screening and assessments
Of hiring managers attribute bad hires to the pressure of filling a position quickly

Hundreds of candidates, not enough time
Questions & answers only, no AI Insights tab. Per-question 1-5 star ratings, an aggregate "average rating" up top, nothing else to go on.
Assist judgment, don't replace it
Customers were asking for faster decisions at scale, but the harder product problem sat underneath: how much authority should AI have in a hiring decision?
We explored whether AI should verdict the evaluation, sit passively as another source of information, or act as an assistant that gives recruiters useful evidence while keeping the final judgment with them.
The direction I advocated for was the third: AI should make evaluation faster and more informed without becoming the decision-maker.
How do you give evaluators a genuinely useful AI analysis, without teaching them to stop applying their own judgment?
Confidence in the evaluation, not a final verdict
Lock 3 of a job's 10 attributes for the AI to evaluate while configuring the screening. Every completed screening opens on AI Insights by default, showing its confidence, never a yes/no of its own.

3 of 10 attributes, locked
Standard or Conversational screening mode, a questionnaire, and 3 of the job's 10 attributes locked in for the AI to evaluate, set once, applied to every candidate for that role.

Summary before verdict
Who this candidate is, in plain terms, then the candidate analysis: Advancement strengths and Development areas, 3-6 bullets each. Context before anything recommendation-shaped, to reduce anchoring.

Confidence in the read, not a score
Low / Medium / High confidence per locked skill. "Show reasoning" is collapsed by default and expands to the exact source question and quoted answer to provide substantial understanding of the AI reasoning for the skill assessment.

Per-item status, never an aggregate
Each requirement gets its own Met / Partially met / Not met tag with a line of evidence, no "4 of 5" fraction anywhere, so it can't be skimmed as a pass/fail verdict.
Volume outpaced judgment
Customers were asking for faster decisions on high-volume roles, where evaluators had the least time to give each candidate a fair, in-depth read. Alongside my UX design colleague, I analyzed competing products that were already using AI to transcribe, summarize, and evaluate candidate interviews.
We used that research to define our recruiter persona, identify where existing approaches placed AI on the spectrum between information and decision-making, and determine how our feature should be positioned: useful behavioral insights that support a recruiter's judgment rather than replace it.
That tracked with the Challenge: this was built to help evaluators decide, not to be a sales pitch.
The number of candidates to evaluate kept growing while evaluator time didn't. That tradeoff meant either slow, late evaluations or fast ones that missed what the candidate actually said.
The goal wasn't to replace human judgment, but to support it with objective, evidence-backed insights that removes subjective guesswork and gives evaluators a reason to trust their own call.
Where this sits against the market
Sapia and Workable score and rank candidates with less traceability behind the number.
HireVue, BrightHire, and Metaview surface evidence but never render a verdict at all.
The shipped direction sits in the narrow band that does both, moderate authority, fully traceable. A spot none of the five actually occupy, since most pick one extreme.
Note: The graph is based on how I read each company's own public product pages, my interpretation of their positioning, not a professional market or product audit.
Before asking engineering to build this, we needed to know the AI could actually analyze answers against criteria reliably, not just assume it, guess at it, or hope.
Cursor for the prototype, n8n orchestrating the pipeline end to end, Supabase and CloudConvert handling storage and conversion, OpenAI transcribing and analyzing the answers.
Running it myself surfaced where it breaks: weak prompts, hallucination past a certain number of criteria, vague answers throwing it off. That shaped what we asked engineering to build.
Four real versions, before landing on the one that shipped
No formal A/B test. Real prototype iterations, each one testing a distinct failure mode rather than a small tweak on the last.

Balanced evaluation
ShippedSummary, advancement/development, core skills confidence, requirements

Too much AI authority
1-10 score per criterion, pros/cons, overall AI verdict

Evaluator bias
Evaluator picks 3 of 10 criteria, generates on demand delivering inconsistency

Ungrounded score
Per-question notes only, plus a candidate score with no visible reasoning behind it
Each rejected version fails for exactly one reason, not a vague “we liked A better”:
Four personas: One weak, one strong, two in between, tested against real AI output across all four versions.
I built a fuller closed-environment prototype with Cursor: a working replica of the product with a questionnaire builder, screening links, and live database. This gave us a realistic environment to run the comparison with real AI output rather than static mockups.
Each teammate submitted their persona's answers through a real screening link, and the AI running on the OpenAI API's latest model at the time, transcribed and analyzed each submission automatically, across all four UI versions.
The comparison wasn't about which layout we preferred. I wanted to know whether the AI's read actually tracked candidate quality consistently, using real output rather than mockups. Building and running the prototype took a few days and gave us evidence before committing engineering effort to the wrong interaction model.
Three real forks
Hiring decisions were never tied to one answer. The AI needed to tell one whole-candidate story, not score answers in isolation, which risked bias and didn't map to how the decision actually gets made.
A separate profile-level fit score already existed. A single interview is only one of several stages, scoring off one interview risked losing candidates who'd do better later, and handing evaluators a subjective, biasing number.
Confidence per skill is calibrated mainly from behavioral data, decision patterns and downstream candidate progression, learned continuously without requiring anyone to click anything. Explicit thumbs up/down stays optional, with a light prompt only on thumbs-down. Required feedback was the right instinct pointed at the wrong mechanism, coverage matters more than getting everyone to rate.
Reasons to advance, areas for development, and per-skill confidence already say what's needed. A separate AI verdict next to the evaluator's own Yes/Maybe/No stops assisting and starts voting, least useful on obvious candidates, riskiest on ambiguous ones.
An early version scored each skill 1-10, dropped for putting the AI in the position of judging the candidate. Reframed to measure the AI's own confidence instead, same visual weight, a very different claim.
An earlier version let evaluators pick 3 criteria themselves, per candidate, on demand. Testing surfaced it fast: different evaluators picked different criteria for the same role, and could indirectly steer the AI toward a candidate they already liked. Locking the 3 criteria once, at setup, fixed both problems.
A focused product isn't a limited one, it's a positioned one
The confidence score had to be built up from evaluator feedback and progression data over time.
Two designers, one PM, engineering, shaped what got explored vs. shipped inside this phase.
The real risk isn't a shipped bug, it's evaluators leaning on the score instead of their own judgment, especially early on.
Speed is easy, earned trust is not
Building the n8n prototype changed how I think about the role of a designer in AI product development. A working prototype let me test the AI end to end, expose failure modes, and influence product decisions before engineering committed to implementation.
AI can genuinely help evaluators move faster. The open question this feature left me with is whether that speed comes at the cost of blind trust, especially at a stage that feels low-stakes.
Good applied AI looks less like a landing page generated in ten minutes and more like a closed environment built to actually test a feature first.