AI
Ordinary Video Is Aimed at a Bedside Apraxia Score
A 2026 hypothesis paper says ordinary video could score ideomotor limb apraxia for trials and telemedicine, if it beats injured controls without apraxia.
A hypothesis paper dated August 18, 2026 argues that ordinary video can score ideomotor limb apraxia as a measurement adjunct.
Neurologist Alireza Minagar set out a camera pipeline, a control-group design, and the results that would kill the claim. No patients were recruited.
The Paper Stops Short of a Diagnosis
Ideomotor limb apraxia is a higher-order failure of learned or purposeful movement. Weakness, sensory loss, ataxia, extrapyramidal disease, poor comprehension, and lack of cooperation do not adequately explain it. It is common after left hemisphere stroke and shows up in several neurodegenerative diseases. Hugo Liepmann framed the modern idea more than a century ago. Clinics still catch it by watching people gesture.
Minagar’s proposed framework and validation design treats a markerless camera system as a measurement adjunct a clinician can override. The paper does not offer an autonomous diagnosis. It states two claims that can fail separately. Continuous spatial and temporal features taken from ordinary video should distinguish apraxic from non-apraxic gesture and should track severity. A model trained on blinded expert ratings should approximate those ratings on new patients.
The second claim is capped on purpose. A model trained on experts can, at best, reproduce those experts. Agreement with a panel does not prove the model measures the disorder more accurately than the panel does. Expert consensus is a reference standard for criterion validation, not ground truth. Functional tool use, lesion maps, diagnosis, and change over time would have to carry the measurement claim as well.
The critical comparison is therefore not apraxia versus healthy controls, but apraxic patients versus neurologically affected patients with comparable weakness, lesion burden, age, and general cognitive impairment but without apraxia.
Alireza Minagar, neurologist, hypothesis paper dated August 18, 2026
The pipeline has two analytical levels. Level one scores how a gesture was executed, in space and in time. Level two asks whether the observed action matched the requested target. Kinematics can carry the first level. Content errors, substitutions, perseveration, and omissions sit mostly at level two, and the paper treats them as an exploratory secondary endpoint. Body-part-as-object errors, in which the hand stands in for a tool, are kept out of the first training vector because they depend on fine finger shape and hand orientation in depth, the quantities single-camera pose recovers least reliably.
Lab Kinematics Never Reached the Ward
Quantitative study of apraxic movement is old. What never followed was a method a ward or a telemedicine clinic could run without a motion-capture lab. That is the gap the 2026 paper tries to close with commodity video.
FROM LAB GRAPHICS TO A BEDSIDE CAMERA
- 1990: Howard Poizner, Loraine Mack, Mieke Verfaellie, Leslie Gonzalez Rothi, and Kenneth Heilman publish three-dimensional computergraphic analysis of apraxic gestures in Brain.
- 1996: Joachim Hermsdorfer and colleagues show that apraxic imitation trajectories differ from those of controls on kinematic analysis.
- 1999: Kathleen Haaland, Deborah Harrington, and Robert Knight report spatial deficits in aiming that are specific to ideomotor limb apraxia, using left-hemisphere-damaged patients without apraxia as the comparison group.
- 2010: Tim Vanbellingen and colleagues publish the Test of Upper Limb Apraxia, a 48-item battery with a six-point score across imitation and pantomime for non-symbolic, intransitive, and transitive gestures.
- 2011: The same group publishes the Apraxia Screen of TULIA, a 12-item bedside apraxia screen derived from that battery.
- 2020: Google’s MediaPipe Hands paper describes on-device tracking of 21 hand landmarks from a single RGB frame.
- 2024: Justin Huber, Stacey Slone, and Jihye Bae show that modest cameras can recover recommended upper-limb kinematic metrics in a drinking-task pilot.
- August 18, 2026: Minagar publishes the hypothesis and the validation design, with no patient recordings attached.
TULIA improved reliability in stroke patients and healthy controls. Scoring still rests on a human watching the movement, live or on video. Agreement is stronger for summary scores than for individual items. The output is an ordinal label, not a continuous measure of path length, submovement count, or endpoint error. A clinician can write that a salute was moderately impaired. The note does not say how far the hand left the target path.
Among patients with a first left hemisphere stroke, M. Donkervoort and colleagues found apraxia in 28% of 338 people treated in rehabilitation centres (96 of 338) and in 37% of 154 people in nursing homes (57 of 154). Those shares are large enough that a weak, observer-only score is a practical problem for trials and for follow-up, not a rare bedside curiosity.
Markerless Stroke Tools Still Target Paresis
Open-source pose estimators now recover joints from a webcam. OpenPose uses part affinity fields for multi-person pose. MediaPipe Pose tracks 33 body landmarks. MediaPipe Hands tracks 21 landmarks per hand. The Holistic landmarker combines body, face, and hand landmarks into 553 points from stills, decoded video, or a live feed. Transformer pose models such as ViTPose sit in the same toolbox. Stroke research has already pointed those tools at paresis, range of motion, compensation during reaching, and automated Fugl-Meyer-style scoring.
Huber’s group, writing in Scientific Reports in 2024, recovered kinematic metrics of the drinking task with ordinary cameras and an open-source pose model, then compared them with marker-based motion capture in 10 neurotypical adults. Mean error was -0.12 movement units, 3.4 mm of trunk displacement, and 0.15 seconds of movement time, with no Bland-Altman bias. That is evidence that commodity video can carry the kinematic variables stroke labs already care about. It is not evidence about praxis. The drinking task measures a reach-and-sip. Apraxia lives in how a learned or novel gesture is planned and shaped, including pantomimes that hide the hand.
A model that only showed machine learning can parse an upper limb after stroke would be derivative. Minagar’s narrower claim is the intersection that the stroke-vision literature has not occupied: markerless video, mapped onto the spatial and temporal error dimensions of an established apraxia scale, checked against blinded expert consensus, and built for bedside or remote use. The paper asks for a formal literature search before anyone treats that gap as proven absence.
Why Healthy Controls Would Inflate Accuracy
If apraxic patients were compared only with healthy adults, a classifier could look sharp by detecting a slow, short, clumsy reach. That is left hemisphere stroke with a weak arm, not apraxia. Haaland’s 1999 aiming study already used left-hemisphere-damaged patients without apraxia for that reason. Minagar makes the same comparison the primary test. Healthy controls may sit in the design as a reference. They cannot be the main contrast.
Group B is defined by the absence of apraxia, not by the absence of abnormal movement. Spasticity, intention tremor, and cerebellar ataxia all fragment velocity profiles, add extra peaks, and bend paths, which are the same surface features proposed as apraxic markers. At the level of one trial those signatures overlap. The paper’s discriminating principle is content dependence. Pyramidal and cerebellar degradation should scale with mechanical demand and stay largely indifferent to whether the act is a meaningful transitive gesture, a meaningless posture, or the same action done with the real object in hand. Apraxic degradation should dissociate across those conditions, with pantomime worse than actual tool use and command worse than imitation.
The primary quantity is therefore a within-person contrast across elicitation conditions, for example the difference in trajectory deviation between pantomimed and object-mediated versions of the same action. Every participant in both groups would have to perform matched actions under each condition. Three secondary contrasts are stated as hypotheses, not as facts in hand. Intention tremor occupies a comparatively narrow band around 3 to 5 Hz and should be removable by spectral decomposition before submovement counting. Spastic movement should cut peak velocity while preserving trajectory shape. Cerebellar dysmetria should show up as terminal overshoot or undershoot, while apraxic spatial error should appear earlier as a wrong plane or a wrong hand orientation.
HOW THE THREE SCORES DIFFER
| Score | Items and scale | What the clinician gets | What it takes to run |
|---|---|---|---|
| TULIA | 48 gestures, six-point ordinal | Summary reliability with item-level scatter | A trained rater; video rating is treated as the gold standard even in clinic |
| AST | 12 gestures, pass or fail, total 0 to 12 | Cut-offs of 9 (mild) and 5 (severe); Cronbach alpha 0.92; test-retest ICC 0.95; stroke PPV 100%, NPV 92% | About 3 minutes at the bedside, no extra hardware |
| Proposed video adjunct | TULIA-aligned tasks, continuous features | Spatial and temporal profile per trial, plus a limited detection and grading signal | Ordinary RGB video at 60 fps, expert labels for training, optional second camera |
The AST was cut from TULIA on 133 stroke patients and 50 healthy controls, then checked in a new cohort of 31 stroke patients. Shirley Ryan AbilityLab’s instrument review notes that aphasia with severe comprehension problems can drag pantomime items below cut-off for the wrong reason. A camera pipeline inherits that language confound unless the protocol separates imitation from command, which the paper requires.
Pantomime Puts the Hand Over the Face
Single-camera pose estimation degrades under self-occlusion. Apraxia pantomime is full of those movements. Combing the hair, brushing the teeth, and saluting bring the hand across the body or in front of the face. Fine finger configuration and hand orientation in depth are the principal technical vulnerability, and the paper says they must be validated against manual annotation or motion capture in a subset rather than assumed.
A two-camera setup is the default for the validation study: one frontal view and one oblique view at about 45 degrees, same height and distance, with a hardware trigger or a visible and audible marker for sync. That arrangement resolves most midline and face occlusions and lets the pipeline either triangulate or pick the view with higher per-keypoint confidence. The single-camera path stays in the design as the minimum bedside and telemedicine configuration, because that is the configuration the deployability claim rests on. The gap in keypoint stability and in downstream classification between one view and two views is a prespecified secondary comparison.
At 30 frames per second the highest frequency that can be represented without aliasing is 15 Hz, which is adequate for duration, onset latency, and well-separated submovements, and marginal for closely spaced velocity peaks. Features that use the second or third derivative of position are unreliable at that rate because numerical differentiation amplifies high-frequency noise. The paper therefore recommends 60 frames per second for the validation study, with every session’s frame rate logged and used as a covariate in any multisite analysis. Pooled recordings would be resampled to a common rate and filtered with the same low-pass cutoff so that a difference in submovement count is about the movement, not about the camera.
RECORDING RULES THE VALIDATION WOULD NEED
- Light and background: At least 300 lux at the participant, frontal and diffuse, with no window in frame, and a plain matte backdrop that contrasts with clothing and skin.
- Clothing and jewelry: Short sleeves or close-fitting cloth to above the elbow, watches, bracelets, and rings off.
- Camera placement: Fixed tripod, lens at seated shoulder height, perpendicular to the frontal plane, 1.5 to 2.0 metres away, head through hip and both arms in frame for the full range.
- Image settings: Minimum 1920 by 1080 at 60 frames per second, fixed focus, fixed exposure, automatic white balance off.
- Chair and calibration: Armless chair of fixed height, feet flat, back unsupported, plus a 10-second rest-and-shoulder-abduction calibration at the start of each session.
Pose estimator family is an experimental factor, not a settled choice. A lightweight convolutional model such as MediaPipe or BlazePose would run beside a transformer model such as ViTPose on identical recordings. Benchmark wins on natural-image datasets do not transfer automatically to seated hemiparetic patients, self-occlusion during pantomime, uncontrolled lighting, and modest resolution. Larger transformer variants may also miss the commodity-hardware claim. If the two families prove interchangeable for the features that matter clinically, that result is useful.
Inertial units on the upper arm, forearm, and hand are an optional research arm, not part of the deployable system. Vision supplies an absolute spatial frame and then fails under occlusion and poor light. Inertial sensing is continuous and indifferent to occlusion, then drifts and has no absolute position. An extended Kalman filter could hold segment pose in the state, propagate with a rigid-body model, and update with vision landmarks weighted by the pose estimator’s own confidence, so occluded points are downweighted automatically. Requiring sensors would forfeit the bedside and telemedicine advantage. The fused arm is there to set an upper bound: if fusion substantially beats video alone, the shortfall is measurement quality; if it does not, the video-only claim is stronger.
Soft Labels Cap What the Model Can Learn
Apraxia cohorts are small, likely fewer than 100 patients, while each video is high-dimensional. The paper therefore prefers interpretable classifiers such as random forests or support vector machines over deep networks trained from scratch. A transformer over feature sequences is a prespecified secondary arm, admissible only with pretraining or transfer, strong regularization, and strictly participant-level evaluation. At realistic sample sizes the simpler model may win, and that result would be reported as such.
Ground truth would come from at least three trained raters using a prespecified manual and a validated instrument, with interrater reliability reported before any consensus. Disagreement is retained. Forcing a majority label discards information on the trials that are clinically ambiguous. Soft-label training presents each trial as the distribution of panel ratings, so a trial called impaired by two of three raters is a target of about two-thirds. The objective is cross-entropy against that distribution. Trials with disagreement stay in training and in evaluation. Panel reliability, as an intraclass correlation for severity and a chance-corrected statistic for error categories, is the realistic ceiling. The model is not expected to beat the raters who labeled it.
Evaluation is leave-one-subject-out. Splits occur at the participant level, never at the trial level, so no frames from a test patient leak into training. Feature selection and tuning stay inside folds. Primary metrics are the intraclass correlation between predicted severity and expert consensus, and, for binary detection, the area under the ROC curve. Generalization claims are ordered by difficulty: new trials of trained gestures, new patients, unseen gestures, different elicitation conditions, and a different camera or site. Unseen gestures and unseen sites are a much stronger claim than new patients doing the same standardized set.
THE NUMBERS THE PROPOSAL IS BUILT ON
- Capture rate: 60 frames per second at 1920 by 1080, logged per session, with 30 fps treated as adequate only for duration and well-separated submovements.
- Working distance: 1.5 to 2.0 metres, lens at seated shoulder height, optional second camera at about 45 degrees.
- Light floor: 300 lux at the participant, no window in the field of view.
- Tremor band: 3 to 5 Hz reserved as a covariate for intention tremor, not as an apraxia marker.
Candidate level-one features are movement onset latency, duration, peak velocity, velocity-peak count, path length, trajectory deviation, endpoint error, gross joint-angle relationships, pauses and corrections, intersegmental coordination, and trial-to-trial variability. Spatial deviation can be a dynamic time warping distance to a normative template. Temporal disruption can be velocity-profile irregularity and fragmentation. The emphasis on hand-crafted, clinically motivated features is deliberate. Training a deep network from scratch on a few dozen patients would invite overfitting.
If the hypothesis holds, those continuous features would differ between apraxic patients and the appropriate comparison group, reproducing at the patient level the group differences laboratory kinematics have already shown. Severity estimates would correlate with expert consensus on the reference battery. Error-dimension profiles would align with clinician-assigned types for the spatial and temporal dimensions, with weaker performance expected on content errors. Refutation is specified in advance. If computed features fail to separate apraxic from control performances once weakness and general brain damage are controlled, or if a model trained on expert ratings fails to generalize to new patients, new raters, or new recording conditions, the stronger form of the hypothesis is done.
Consent Tightens When Video Is Identifiable
Apraxia is still easy to miss beside hemiparesis and aphasia. A score that only detects a slow, short reach would be a false win, which is why the paper spends so much of its length on Group B and on within-person contrasts rather than on model architecture. Gesture exams in clinic still turn on pantomime, imitation, and errors such as using the limb as if it were the tool. Those are exactly the motions that hide the hand from a single camera. The technical risk and the clinical risk are the same risk.
Ethical development would need consent for video capture and storage, protection of identifiable recordings, and a rule that an automated score is not granted more authority than its validation supports. Minagar states three specific positions. When the clinician and the automated output disagree, clinical judgment prevails without a required justification, both scores are kept, and systematic disagreement is audited. When an automated score is wrong and contributes to a mistaken assessment, responsibility stays with the assessing clinician, and validation limits must sit beside every output. In a population where aphasia and cognitive impairment are common, consent cannot rest on written or verbal comprehension alone. Simplified visual and demonstrative formats, a legally authorized representative where capacity is absent, revocable consent including during a recording, and re-consent if capacity changes are the minimum. Video is identifiable in a way that an ordinal TULIA total is not, which raises the threshold for all three safeguards.
Should even the bounded form hold, severity would become comparable across examiners, institutions, and time, which matters for individual monitoring and for trials that lack a reliable quantitative endpoint. Ordinary video could extend structured assessment to clinics without a behavioral neurologist, including remote links, and could give therapists a measurable trace of gesture production. None of that removes the clinician. If the within-person contrasts fail to separate Group A from Group B once weakness and lesion burden are accounted for, the stronger form of the hypothesis fails with them.
Disclaimer: This article is news reporting and analysis of a published scientific hypothesis and related measurement studies. It is informational only and is not medical advice, a diagnostic method, or a recommendation to examine or treat any patient with a camera system. Readers should consult a qualified neurologist, physiatrist, or occupational therapist before changing how apraxia is assessed or how rehabilitation is planned. Figures, cut-offs, and design details reflect the cited paper and instrument sources as published, and they may change if a validation study is run or if scoring manuals are revised.
-
AI3 months agoFable 5 Came Back Under a Commerce On-Off Switch
-
AI4 months agoGoogle’s SpaceX GPU Lease Has a Sept. 30 Deadline
-
CRYPTO4 months agoPlasma One’s XPL Locks Face a 1.81 Billion Cliff
-
APPS4 months agoDGO’s Rs 549 World Cup Pass Cost Fans Sleep and Data
-
AI4 months agoMoonshot AI’s $30 Billion Ask Became a $35 Billion Close
-
NEWS4 months agoColorOS 17 Device List Spans Oppo, OnePlus and Realme
-
GAMING4 months agoXbox Cuts 3,200 Jobs After Five Years of Thin Returns
-
GAMING3 months agoThe RTX 4050 Under Rs 70,000 Hides a Wattage Gap
