Skip to content

Method

For these video self-reports, interpreting timing differences, deploying a personal correction, and keeping the profile for later require separate reviews. Researchers assemble the evidence; the program returns a next step and records its reason, the responsible person, and review conditions.

People judge the evidence and remain responsible. The program does not check whether a study is true or complete.

Protocol overview: proposed use, human evidence review, ordered rules, actions, and a decision record with owner and review trigger.
Fig. 1. New evidence or a changed use can reopen the record for review. Open full-size figure ↗

Review Rules

R1Scope of Use

Does the conclusion reach beyond who, where, or what was tested?

A result may depend on the people, setting, or information tested. If broader use is unsupported, limit its interpretation and use to that tested scope.

View formal conditions and code
R2Added Benefit

Do added sensors or automatic changes really improve the result?

Compare against the alternative already tested for this task. When added benefit has not been demonstrated, keep that tested comparator.

View formal conditions and code
R3User Outcomes

Did a metric improve, or were user outcomes also evaluated in use?

A better measurement or stated preference does not establish a benefit in actual use. If that is the only evidence for deployment, first run a preregistered study with users in context.

View formal conditions and code
R4Long-term Reuse

Has long-term reuse been checked for stability, transfer, and population coverage?

A profile may not remain valid in another session or setting, and its reference data may not cover the relevant people. Missing or incomplete evidence means holding off on long-term storage and reuse elsewhere.

View formal conditions and code
R5Permission Conditions

Is there direct evidence for this use, with user control or reversibility?

Reversible personalization also needs direct user-outcome evidence that matches the proposed use. Only when all permission conditions are met and no other rule requires a restriction can it proceed within tested limits.

View formal conditions and code

Insufficient evidence never means approval by default.

Execution order and permission boundaries

R4 → R2 → R3 → R1 → R5

The questions above are numbered for explanation, not execution. Rules run in ascending priority order; the first match determines the result. The numeric priorities are 10, 20, 30, 40, and 90, respectively.

  • R1 limits the scope of a result; it does not authorize personalization.
  • R5 must independently satisfy all its permission conditions, with no higher-priority constraint matched.
  • When no rule matches, the program requests the specific missing evidence; it never approves by default.
Inspect the resolver
Calibration Example

A correction learned from separate calibration videos reduced error against a reference trace. Does that justify deployment?

  1. Record the evidence: comparator improvement is supported, but the evaluation measures reference error; direct user benefit was not tested.
  2. Resolve the rules: R4 and R2 do not match this record. The proxy evaluation activates R3 before R1, requiring a preregistered in-context user study.
  3. Reopen with direct, use-matched evidence and prespecified benefit and harm criteria. Re-run all rules, including the remaining R5 conditions.
View the existing calibration record

Research Questions

The same report can describe this session, inform a personal timing correction, or become a profile for later use. Being able to compute these results does not establish that each use is justified.

We need separate evidence for what the differences mean, whether additional information helps, and whether correction benefits users. These decisions motivate our three research questions.

RQ1
Can the same rules review different uses and return the expected results?
RQ2
What can a profile tell us when measurement conditions change?
RQ3
What useful evidence do sensors or feedback add?