
From Signal to Support · AI research case study
The strongest AI decision was rejecting the model.
In 2020, our four-person hackathon team explored passive facial analysis as a way to connect students with support. Further research exposed a critical inference gap—including evidence we should have found then. I rejected the mechanism, preserved the valid human problem, and evaluated whether a consent-first alternative still deserved to exist.
See why the model was rejected- Historical artifact
- Verified
- External evidence
- Sourced
- Passive mechanism
- Rejected
- Consent-first alternative
- Proposed
Executive finding
A model can classify a face and still fail the decision.
The original concept treated three different questions as though they were one: whether software could classify visible facial movements, whether those movements reliably revealed internal emotion, and whether that inference justified a support decision.
The first is about technical buildability. The second is about construct validity. The third is about whether an output is safe and useful in a consequential workflow. Further research did not support the leap from the first question to the next two.
Decision Reject passive emotion inference as the mechanism.
The 2020 hypothesis
What our team proposed—and what we did not prove.
Software would run alongside remote learning, observe facial movements over time, identify departures from an individual baseline, and use those changes to prompt a connection to support.
Can a system label visible patterns?
Yes, in a bounded benchmark sense. AffectNet contains more than 450,000 manually annotated images. Yet categorical agreement in its double-annotated subset was 60.7%. The labels represent observer judgments—not ground truth for private emotion.02
Does the label measure internal emotion?
Not reliably enough for the proposed claim. Facial movements vary across situations, cultures, people, and emotion categories; context materially changes interpretation.010304
Should the output trigger intervention?
No evidence established that it should. A confident expression label cannot explain what a student feels, why they feel it, whether support is wanted, or whether an intervention would help rather than harm.
Improving a classifier would not repair an unsupported inference.
Evidence timeline
The scientific problem did not appear years later.
The updated evidence matters, but the most important correction is historical: a major review challenging the inference was already available before our original work.
A benchmark is built
AffectNet made large-scale expression classification more buildable. Its labels came from human annotators, however—not direct access to anyone’s internal state. 02
The inference is challenged
A major review found facial configurations and emotion categories have variable, context-dependent, many-to-many relationships. This evidence existed before our presentation. 01
Our team proposes passive sensing
The hackathon concept treated changes in facial expression as evidence of a student’s emotional state and possible need for support. It remained a concept. 00
A major vendor narrows use
Microsoft retired general-purpose emotion-inference access, citing scientific consensus, generalization, and privacy concerns. 06
Expression is not inner state.
Amazon’s own documentation makes that distinction. Microsoft’s retirement of open-ended emotion inference adds a product-policy signal—but is not independent proof.0506
Education is a restricted context.
Article 5(1)(f) of the EU AI Act prohibits covered biometric emotion inference in educational institutions, with narrow medical or safety exceptions. This is an EU rule, not a global ban.07
Validity claims create exposure.
The FTC flags biometric privacy, security, bias, and misleading-claim risks; its statement is nonbinding. NIST offers a voluntary risk framework, not a certification.0809
The problem that survived
Students do not need to be watched to need help navigating support.
The evidence supports an access and navigation problem without supporting passive emotion inference as its solution.
of 61,085 community-college survey respondents disagreed or strongly disagreed that they would know where to seek professional mental-health help.
CCCSE report + methodologyof respondents who reported needing help during the prior year said they never sought it. This does not establish that navigation was the sole cause.
CCCSE report + methodologyof Healthy Minds respondents reporting no or fewer services selected “not sure where to go.” The denominator is this subgroup—not all surveyed students.
Healthy Minds reportThese large participating-institution surveys establish navigation uncertainty; they do not establish demand for this product, causal impact, or a nationally representative rate for every U.S. student.
Help students state what they need, understand their options, and reach a human who can help—without attempting to infer their private emotional state.
Consent-first product hypothesis
Ask. Clarify. Offer choices. Let the student decide.
The redesign replaces hidden inference with a visible, student-controlled workflow. It is a proposed design direction—not a diagnosis system, crisis service, or validated intervention.
- 01Ask
The student chooses whether to request help.
- 02Clarify
Bounded questions establish the job, urgency, and constraints.
- 03Offer
Verified options explain what each service can and cannot do.
- 04Choose
The student selects an option and controls what is shared.
- 05Connect
A trained person owns judgment and the handoff.
- 06Confirm
Follow-up checks whether support was actually reached.
Explicit exclusions
Trust is product functionality.
- No passive camera or microphone
- No diagnosis or crisis prediction
- No autonomous clinical referral
- No invisible data sharing
- No secondary use without fresh consent
Startup strategy
A real problem does not automatically make a defensible startup.
Market sizing should expose assumptions, not hide them inside a large dollar headline. The honest order is TAM → SAM → SOM, and the current evidence does not support a revenue forecast.
Total addressable market
Context: 16.4 million U.S. undergraduates at degree-granting Title IV institutions in fall 2024. That is a population, not revenue.12
eligible institutions × buyer-validated annual contract valueServiceable available market
Starting segment: 1,015 AACC-defined community colleges, serving 6.5 million credit and 4.1 million noncredit headcount—before buyer, workflow, privacy, capacity, and incumbent-vendor filters.13
qualified segment − operational and competitive exclusionsServiceable obtainable market
The first honest milestone is one paid institutional pilot followed by a verified expansion. An arbitrary percentage of enrollment would create precision without evidence.
qualified accounts × paid conversion × retained ACVCompetitive reality
The market converged on navigation while weakening the novelty claim.
| Alternative | Existing capability | Implication |
|---|---|---|
| TimelyCare 14 | Centralized intake, matching, follow-up, human navigation, warm handoffs, reporting | Direct feature-level overlap validates demand and compresses novelty. |
| Findhelp 15 | Configurable resource networks, screening, referrals, follow-up, outcomes | A new entrant cannot win by building another directory. |
| Campus process | Advisors, counseling, basic-needs offices, portals, CRMs, manual case management | “Use what we have” is a serious competitor in a budget-constrained segment. |
Validation gates
What would have to be true before software is built.
Campuses show measurable drop-off between requesting help and receiving it.
Stop if capacity—not navigation—is the binding constraint.A named owner funds a bounded evidence pilot from a real operating or grant budget.
Stop if interest only produces unpaid pilots.Buyers can state why this beats TimelyCare, Findhelp, or their existing process.
Stop if the answer is only “it uses AI.”Students understand the data flow and participate voluntarily; campus owners approve the boundary.
Stop if trust depends on hidden use or mandatory participation.The workflow improves completed, appropriate connections and time to human contact.
Stop if it only increases clicks or form completions.Cost per completed connection and staff time compare favorably with the alternative.
Stop if human navigation makes delivery structurally uneconomic.Primary success measureA student who explicitly asks for support reaches an appropriate, verified resource or human navigator within the agreed service window.
Decision memo
Stop the mechanism. Preserve the problem. Hold the startup thesis.
- Passive facial-expression inference
- Rejected
- Student-support navigation problem
- Supported
- Consent-first workflow
- Proposed
- Product validity
- Not established
- Standalone startup thesis
- Hold
- Research case study
- Publish
What this demonstrates
Research is not defending the first answer.
My contribution is not a claim that the original concept was right. It is the decision process used to correct it.
- 01Construct validity
Separate what a model predicts from what the product claims that prediction means.
- 02Dataset scrutiny
Read annotation methods and disagreement before accepting a benchmark as ground truth.
- 03Evidence synthesis
Trace consequential claims to peer-reviewed or official sources and preserve disagreement.
- 04Responsible AI
Evaluate consent, privacy, governance, failure modes, and human accountability before deployment.
- 05Product reframing
Preserve a valid user problem while replacing an invalid technical mechanism.
- 06Market falsification
Test buyer, workflow, competition, and moat before turning a population into a forecast.
That is the standard I bring to AI research and startup strategy: identify what can be built, determine what should be built, and preserve the distinction.
Source register
Every material claim has a visible evidence trail.
Source types are identified so a product page, voluntary framework, policy statement, peer-reviewed study, and binding law are not treated as interchangeable.
- 00
Primary artifact · Private
Original 2020 Edu-nomics team presentation
Supports the historical concept, four-person team scope, proposed passive mechanism, and absence of build or validation evidence. Reviewed privately and not republished because it contains teammate photographs and third-party template assets.
- 01
Peer-reviewed review
Barrett et al. — Emotional Expressions Reconsidered (opens in new tab)
Supports the scientific limits of inferring specific emotion categories from facial movements across people and contexts.
- 02
Peer-reviewed dataset paper
Mollahosseini et al. — AffectNet (opens in new tab)
Supports the scale and human-annotation methodology of a major facial-expression benchmark, including reported annotator agreement.
- 03
Peer-reviewed study
Gendron et al. — Emotion perception across cultures (opens in new tab)
Supports cultural and contextual variation in how facial movements are interpreted.
- 04
Peer-reviewed study
Goel et al. — Context is critical for emotion perception (opens in new tab)
Supports the material role of context in human emotion perception.
- 05
Official product documentation
Amazon Rekognition — Emotion API (opens in new tab)
Explicitly distinguishes predictions based on facial appearance from a determination of internal emotional state.
- 06
Official vendor policy
Microsoft — Responsible AI framework and restricted capabilities (opens in new tab)
Supports Microsoft’s retirement of general-purpose emotion-inference access. This is a vendor policy signal, not independent scientific proof.
- 07
Primary law
Regulation (EU) 2024/1689 — Artificial Intelligence Act (opens in new tab)
Article 5(1)(f) sets the EU prohibition relevant to emotion inference in education, with narrow medical or safety exceptions.
- 08
Official policy statement
U.S. FTC — Biometric Information and Section 5 (opens in new tab)
Flags privacy, security, bias, and misleading-validity risks. It is a policy statement, not a binding rule by itself.
- 09
Voluntary risk framework
NIST — AI Risk Management Framework 1.0 (opens in new tab)
Supports a risk-management approach centered on validity, reliability, privacy, bias, transparency, and human oversight.
- 10
Student survey report + methodology
CCCSE — Supporting Minds, Supporting Learners (opens in new tab)
Supports community-college navigation and help-seeking findings, with sample sizes and participating-institution context.
Read the methodology (opens in new tab) - 11
Student survey report
Healthy Minds Study — 2024–2025 national data report (opens in new tab)
Supports the narrower “not sure where to go” result among respondents reporting no or fewer services, not all surveyed students.
- 12
Official federal statistics
NCES — Fall 2024 enrollment at degree-granting institutions (opens in new tab)
Supports the 16.4 million undergraduate market-context figure.
- 13
Official sector statistics
AACC — Fast Facts 2026 (opens in new tab)
Supports the 1,015 community-college, 6.5 million credit-student, and 4.1 million noncredit-headcount context.
- 14
Official product page
TimelyCare — Basic Needs Support (opens in new tab)
Supports direct competitive overlap in centralized intake, matching, follow-up, human navigation, handoffs, and reporting.
- 15
Official product page
Findhelp — The Findhelp Network (opens in new tab)
Supports competitive overlap in configurable resource networks, screening, referral tracking, and closed-loop outcomes.
- 16
Official federal guidance
HHS — FERPA and HIPAA guidance (opens in new tab)
Supports the need for case-specific data classification; campus-maintained student health records are often governed by FERPA rather than HIPAA.
Work with me
Is your AI output outrunning its evidence?
I help teams test the link between model output, human judgment, product value, and a commercially meaningful decision.
Evaluate the decision system