Working on the right questions.
I get paid to find the thing that's wrong.
Model outputs, media, arguments — noticing what doesn't fit.
I evaluate AI systems for the ways they fail: synthetic voice and video passing as real, reasoning that holds together right up until it doesn't, answers that are fluent and wrong, outputs that should never have shipped. The domain changes from one week to the next — media authenticity, a technical review, a rubric nobody has stress-tested yet. The habit doesn't: refusing to take a thing part by part, and checking whether the parts add up to the whole they claim.
Both roles are vendor-side, on consumer AI products.