Skip to content

Results

Rule TestsRQ1

Two author-coded records from one menu study led to different actions: keep the tested static menu for the proposed automatic adaptation; allow user customization within the evaluated conditions. Menu case records ↗

Protocol results: five external permissions; deleting R1–R5 changes 1, 2, 2, 1, and 5 actions; structural checks meet their specified responses.
Fig. 3. Actions, changes when rules are removed, and responses to structural tests. Counts refer to author-coded use records. Open full-size figure ↗

The four study records and seven external records cover all five action categories, including five conditional permissions. Removing any one rule changed at least one decision.

These checks show that the program follows its rules. They do not establish that the judgments are correct or that real decisions improve.

Structural checks and source records
ChallengeExpected responseObserved
Order permutationsPreserve outputs100 / 100
Malformed recordsReject input12 / 12
Unrelated-field changesPreserve the focal action5 / 5
Permission-boundary failuresRemove permission30 / 30
Overlapping constraintsFollow declared priority4 / 4

The order tests shuffle graph nodes, edges, artifact-registry entries, rule-list entries, case lists, and external-source entries. Declared priority values remain fixed; the resolver sorts rules by priority before deciding. This is not a test of changing the priority policy.

The permitted uses were a user-adaptable menu, ability-based generated interfaces, a culturally adaptive website, exposed body-sensing calibration, and conservative menu adaptation. Each permission remains inside its source study’s evaluated boundary.

Profile InterpretationRQ2

Estimated timing differences change with the calculation method and reference group.

The reference comes from other people’s self-reports; it is not emotional ground truth.

Measurement figure with profile rank agreement, DTW objective comparisons, shared-reference variance intervals, and signed-calibration gain versus exposure.
Fig. 4. A–C: sensitivity to calculation methods and reference groups. D: calibration gains and annotation time. Variance fractions are descriptive, not reliability coefficients. Open full-size figure ↗

Changing the alignment objective reduced profile-rank agreement to 0.497; changing the reference cohort shifted one profile by up to 0.459 seconds.

Use these differences only to describe the tested conditions, not as an established personal trait. R1 limits interpretation; it does not authorize automatic interface changes.

Additional measurement evidence

Within a trial, timing alignment on one reporting axis helped the other. The axes share a hand, joystick, task, and clock. Across videos, the supported transfer was only a small constant signed shift; it did not establish a persistent personal timing pattern.

Alternative profile definitions had rank correlations of 0.890–0.965 with the consensus definition. Path-mean normalization increased timing magnitude by 0.783 seconds (95% CI 0.570–0.989). Relative and absolute shared-reference variance fractions were 0.760 and 0.714; these describe sensitivity under dependent, overlapping references.

View source

Added InformationRQ3

Do added sensors or user feedback improve on the tested comparator?

Added Sensors

For timing-magnitude prediction without the target-video self-report, the tested sensor stack did not demonstrate improvement over video mean. Errors are in seconds.

MeasureValue
Video-mean MAE0.763 s
Sensor-stack MAE0.794 s
Gain over video mean (95% CI)-0.031 s [-0.081,0.014]

For this task, the added sensors lack demonstrated benefit, so R2 keeps the video-mean comparator. This neither rules out all physiological sensing nor establishes equivalence.

Report Calibration

A correction learned from other videos lowered held-out reference error compared with video-only calibration (Fig. 4D).

Calibration videosGain (label units)
10.021
20.028
40.045
80.063

Better metrics alone do not justify automatic adaptation. R3 calls for a preregistered study in actual use, with direct user outcomes and benefit and harm criteria specified in advance.

Trajectory Analysis

Three additional analyses estimate continuous self-report curves, with errors in label units. They add evidence, not formal review records.

Condition and available informationMatched comparatorMAEGain
Video content · full video available before playbackMatched comparatorMetadata: category, duration, phase
MAE 31.333
MAE30.493Gain0.840
Physiology: content + current and past EEG/fNIRSMatched comparatorFixed content prior
MAE 30.493
MAE30.499Gain-0.006
Post-trial SAM: one rating pair + same-video training reportsMatched comparatorKnown-video population prior, no target feedback
MAE 28.417
MAE26.289Gain2.127
95% CI [0.557,3.878]

MAE and gains are in label units. Lower MAE is better; gain is comparator error minus candidate error. Information timing and validation splits differ across rows, so this is not one ranking.

Three anonymous examples showing observed valence and arousal, metadata and content predictions, the evaluated physiological residual, and post-trial SAM reconstruction.
Fig. 5. Three anonymous trials illustrate estimates from video content, physiological signals, and post-trial feedback. Open full-size figure ↗

Content improved prediction over metadata within the released video library. Adding physiology did not demonstrate a further gain under the primary condition. Post-trial SAM improved reconstruction with feedback, not feedback-free or real-time prediction.

Selection gates and information boundaries

Relative error reduction: 0.069–0.208%. Calculated annotation exposure: 1.7–13.7 minutes, excluding setup and perceived burden.

The primary residual gate requires at least 0.10 label units of improvement on inner-fold predictions. The full evaluated procedure scored 30.499; disabling all residuals instead returns content at every sample, a distinct output. Looser 0.00 and 0.05 gates gave small positive estimates in sensitivity analyses.

The three illustrations and their plotted coordinates have an existing author-confirmed release scope. They exclude complete archives, original identifiers, participant mappings, and stimulus media.

View source

Three anonymous trial illustrations near the 25th, 50th, and 75th percentiles of content-prior error. Every saved one-second sample is shown; examples illustrate the comparisons, not a common ranking.

Review Outcomes

The four uses call for separate decisions. Keeping a profile for later or using it elsewhere still requires evidence of stability, transfer, and reference-population coverage.

Proposed useKey evidence boundaryResolved action
Interpret the profileEvidence boundaryDepends on the measurement and reference conditions.ActionLimit interpretation and use.
Add sensingEvidence boundaryNo demonstrated gain over the matched comparator for this target.ActionKeep the evaluated comparator.
Deploy calibrationEvidence boundaryReference error improved; direct user benefit was not tested.ActionFirst run an in-context user study.
Retain or transferEvidence boundaryEvidence for stability, transfer, and reference coverage is missing.ActionWithhold long-term retention and transfer.

Case Details

What can this result tell us?

Proposed use

Interpret a profile calculated from the continuous self-report collected in one experiment.

Available evidence

The result changes with the interface, calculation method, and reference curve. It has not been checked in a repeat session.

Action and reason

Interpret it only within the tested settings and conditions.

This limits how the result can be interpreted and used. It neither approves personalization nor establishes a stable personal trait.

Evidence for re-review

Evidence matching the proposed interpretation, interface, estimator, reference group, and intended reuse period.

Responsible role in the record

Measurement-analysis lead

Reasoning and formal record

The original measure is a post-trial temporal-deviation profile relative to a reference. R1 limits its scope; permission to personalize is a separate decision.

Original proposal
Compute and interpret a post-trial temporal-deviation profile.
Recorded evidence
The trace is computable, but its magnitude changes with interface, estimator, and shared reference; no repeat session exists.
Official decision
Use the route only inside its evaluated evidence boundary
Source review trigger (verbatim)
Re-review before any durable-profile claim or after a new session, interface, estimator, or cross-video transfer study.
Matched rule
R1 · Interpretation and use cannot exceed the evaluated evidence scopeExecution priority: 40
Fields that triggered the rule
FieldRecorded state
evidence.interpretation_scopePARTIAL
route_capabilityAVAILABLE

These are existing records resolved by the repository program. Re-review conditions summarize outstanding requirements; they are not observed outcomes or promises of permission.

Run a case →

Back to the person watching the video: the tested timing correction reduced error against the reference, but user benefit was not tested. Deploying it therefore requires a preregistered study in actual use; keeping the profile long term requires separate evidence.