Matthew Altobelli

I help teams shipping AI systems catch behavioral blindspots that tank real-world performance

Book a session

AI systems fail in production for reasons the eval never catches: prompts tested only on people who think like the team, surveys that invite the answer they want, rubrics that score fluency when the goal was accuracy. These are measurement and cognition problems, and that seam is where I work.

I am a behavioral scientist (M.S., Experimental Psychology, RIT) who builds. I co-founded Cogbias AI, where I lead R&D on models that detect cognitive bias in written communication, and I was sole architect of a live enterprise psychometric and AI coaching platform: two validated instruments plus the Python, PostgreSQL, React, and multi-agent GPT-4 system behind them. My research on linguistic markers of anxiety and PTSD was presented at Psychonomics and EPA.

I teach psychology at RIT, served eight years in the Air Force, and have led teams of 3 to 100.

Bring me an eval set, survey, prompt, or rubric. You leave with the specific problems named and a plan to fix them.

Rochester NY
AI Consulting & Automation

Services

Survey and Measurement Design Review

For teams collecting human data for training, RLHF, user research, or product decisions. Send the instrument in advance. I review question wording, response scales, ordering effects, and the cognitive biases the design invites, and I return a marked-up version with rewrites

1 hour$250.00
AI Eval and Rubric Audit

"Send me your eval set, rubric, or rater instructions before the call. On the call I walk through the construct validity problems, rater bias exposure, and gaps between what you are measuring and what you are trying to measure. You leave with a written list of specific issues, ranked by how much they threaten your production numbers, and concrete fixes for each

1 hour$250.00
LLM Behavior Diagnosis

Your model or agent is doing something odd with real users and the logs are not explaining it. Bring examples. I apply what we know about human language, expectation, and cognition to narrow down what is happening and what to test next. A short session for a fast read

30 min$125.00