AI evaluation
Compare outputs against a clear rubric, validate claims, surface inconsistencies, and write feedback another person can act on.
Rubric scoring, evidence validation, quality reviewRemote AI evaluation + workflow roles
Hands-on evaluation, UAT, evidence review, and workflow improvement grounded in real operating experience.
Open to fully remote opportunities.

Where I add value
Compare outputs against a clear rubric, validate claims, surface inconsistencies, and write feedback another person can act on.
Rubric scoring, evidence validation, quality reviewTest the real task from a user perspective, document friction precisely, and separate blocking defects from useful improvements.
Workflow checks, edge cases, reproducible findingsMap handoffs, inputs, exceptions, and finish lines so a practical system can hold up outside the demo.
Process clarity, reporting judgment, operational follow-throughEvidence, not adjectives
Different lanes, one operating pattern: understand the work, shape the system, test it, then finish it carefully.
Private prototype
Private prototype. Representative synthetic-data screens are shown here; the source repository is not public.


Active AI media

Live brand


Working fit
I work best where quality, evidence, process, and human judgment have to meet.
If the role needs careful evaluation and operational judgment, the fastest next step is a direct conversation.