Agent guide (headless use)¶
This page is written for a coding agent — or any script — that has to use
Scanpath Studio without a browser: load an eye-tracking-while-reading dataset,
render a figure, read the per-word reading measures your data brings. (AGENTS.md in the repository
root is the opposite document: how to develop this codebase.) A plain-text map
of the whole site is at /llms.txt.
Everything the app draws is reachable from two places:
| Surface | Entry point | Reference |
|---|---|---|
| Python | scanpath_studio.api, re-exported at the package root |
Python API |
| Shell | scanpath-studio render |
CLI reference |
Both go through the same data → measures → plots pipeline as the Streamlit
app, so with the same settings a figure rendered here matches the app's.
Headless use (the API or render) never writes the app's on-device recovery
cache.
The 30-second version¶
import scanpath_studio as sps
words, fixations = sps.load_sample_data() # bundled OneStop demo
pid, tid = sps.list_trials(words, fixations).iloc[0]
fig = sps.plot_scanpath(words, fixations, pid, tid, canvas_size=(2560, 1440))
sps.save_figure(fig, "scanpath.html") # .png/.svg/.pdf need Chrome
canvas_size is the monitor the stimulus was shown on. Pass it whenever you
know it — without it the canvas is estimated from the data extents and the
scanpath no longer sits at its true on-screen position.
Two tables in, one figure out¶
Scanpath Studio works on a pair of tables — word interest areas (one row per
word per trial, with its bounding box) and fixations (one row per fixation).
load_scanpath_data maps whatever your columns are called onto the canonical
fields below — and hands the frames back under your own column names:
fixations["CURRENT_FIX_DURATION"], not fixations["duration_ms"]. Each frame
carries its map in DataFrame.attrs, so every API function reads it, and every
option that names a column (color_by=, word_hover_fields=, …) takes your name
or the canonical one. A column the loader changed (a duration read in seconds,
ids shifted to start at 0) or made (order_in_trial) keeps its canonical name.
The tables below are what load_scanpath_data(..., names="canonical") returns.
A frame merged or concatenated with another table loses attrs: pass the result
through load_scanpath_data again, or work on the canonical frames.
Words (interest areas) — canonical fields:
| Column | Meaning |
|---|---|
participant_id |
Participant id (string). Optional in the source: a stimulus-level word table with no participant column is broadcast onto every trial in the fixations, matched by trial id or else by text_id; load_scanpath_data raises a ValueError when neither matches. |
trial_id |
Trial id; with participant_id it names one trial. Required. |
screen_id, screen_index |
Optional child screen and 1-based order inside a multipart logical trial. Map in both reports. |
text_id |
Which text/passage the row belongs to (plus unique_text_id when the source has a corpus-wide id). |
word_id |
Word index within the trial. Required — it is the join key to fixations. |
text |
The word itself (what gets drawn in the boxes). |
x, y, width, height |
Word bounding box in screen px, origin top-left. Required (or supply left/right/top/bottom, which are converted). |
line_idx |
Source line number, when the export has one. Often constant — the plots derive visual lines from box y instead. |
Pre-aggregated EyeLink IA measures (IA_FIRST_FIXATION_DURATION → first_fixation_ms, …)
and linguistic features (gpt2_surprisal, wordfreq_frequency, universal_pos, …)
are carried through when present.
Fixations — canonical fields:
| Column | Meaning |
|---|---|
participant_id |
Participant id. Optional in the source (a dataset without one becomes a single anonymous participant). |
trial_id |
Must match the words table. Required. |
screen_id, screen_index |
Optional child screen and 1-based order; scientific operations never join across it. |
text_id |
Text/passage id, when present. |
x, y |
Fixation location in screen px. Required unless word_id is given — AOI-sequence data is placed at word-box centers. |
duration_ms |
Fixation duration. Required. |
timestamp_ms |
Fixation onset. Falls back to the row's position within the trial (0, 1, 2, …) when the source has no timestamp — it drives the ordering, so rows must already be in reading order in that case. Those numbers are not times: the internal _timestamp_synthesized column marks them, and the replay lays the fixations end to end by their durations. |
screen_timestamp_ms, screen_fixation_id |
Optional local clock/id that resets per screen; retained alongside the parent-global columns. |
word_id |
Source word/AOI assignment, carried through when the export has one — otherwise NaN. The loader only shifts ids numbered from 1 onto 0-based word boxes; when the export has none, the assignment (box containment, else no word) happens inside the plots that need it. |
order_in_trial |
1-based fixation index, added during normalization. |
fixation_id |
Always present — mapped from the source when it has one, otherwise synthesized as a per-trial running index (1, 2, 3, …). |
saccade_type, saccade_amplitude, eye, pass_index |
Passed through when the source has them. |
Column matching is case- and separator-insensitive: IA_LEFT, ia_left and
Ia Left are the same name.
The minimum a figure needs¶
Boxes, ids, durations. Nothing else — no participant column, no timestamps, no measures:
import pandas as pd
import scanpath_studio as sps
words = pd.DataFrame(
{
"trial_id": ["t1"] * 4,
"word_id": [1, 2, 3, 4],
"text": ["The", "cat", "sat", "down"],
"x": [100, 200, 300, 400],
"y": [100, 100, 100, 100],
"width": [80, 80, 80, 90],
"height": [40, 40, 40, 40],
}
)
fixations = pd.DataFrame(
{
"trial_id": ["t1"] * 3,
"x": [130.0, 320.0, 240.0],
"y": [118.0, 122.0, 115.0],
"duration_ms": [210, 180, 260],
}
)
words, fixations = sps.load_scanpath_data(words, fixations)
fig = sps.plot_scanpath(words, fixations, canvas_size=(800, 300))
sps.save_figure(fig, "minimal.html")
With one trial in the frames, participant / trial can be omitted — more than
one and an underspecified call raises rather than guessing (see
Errors). Neither table had a participant column
here, so both frames come back under one synthetic participant: list_trials returns
participant_id="(all)", trial_id="t1".
Loading real data¶
words, fixations = sps.load_scanpath_data("ia.csv", "fixations.csv")
words, fixations = sps.load_scanpath_data("ia/*.csv", "fix/*.tsv") # globs
words, fixations = sps.load_scanpath_data(words=ia_df, fixations=fix_df)
words, fixations = sps.load_scanpath_data(fixations="fix.parquet") # one table
Ready-made public corpora have their own loaders — sps.load_potec(dir) and
sps.load_onestop(dir) — which return the same normalized pair. See
OneStop.
A raw-gaze table is loaded with load_raw_gaze(path_or_frame) (columns
auto-detected; raw_gaze_schema= overrides), or load_sample_raw_gaze() for
the demo's. Passing it as plot_scanpath(raw_gaze=…) filters it to the trial
and switches the layer on; compare_scanpaths(raw_gaze=…) draws each
trial's samples in its scanpath's color (raw_gaze_b= for a B from another
dataset). It can be the only table: pass None for the words
and fixations (sps.list_trials(raw_gaze=gaze),
sps.plot_scanpath(raw_gaze=gaze, trial="t3")) and the samples are drawn as
recorded. No fixations are detected from them, so animate_scanpath raises
ValueError for a trial without fixations.
For a multipart parent, inspect sps.list_parts(words, fixations, pid, tid),
pass screen="…" to plot_scanpath / animate_scanpath, or call
sps.render_parent_trial(...) for an ordered mapping of per-screen figures.
Data without explicit screen columns can use trial_parts_manifest=; see
Data format.
When auto-detection can't find a column¶
The ValueError names the canonical field, the column names that were tried,
and the columns your table actually has. In full, for a words table whose
columns are subject, para, word, start_x:
Words/IA schema problems: missing Trial ID; missing Word/IA ID; need either (x, y, width, height) or (left, right, top, bottom)
Could not infer these canonical fields from the words/IA table:
- Trial ID (word_schema key 'trial'): no column matched. Looked for: unique_trial_id, trial_id, unique_paragraph_id, paragraph_id, text_id, trial, trial_index, trial_number, presented_stimulus_name, media_name, stimulus
- Word/IA ID (word_schema key 'word_id'): no column matched. Looked for: word_id, IA_ID, ia_index, word_index, aoi, word_idx, char_idx
- Word box (word_schema keys): need either (x, y, width, height) or (left, right, top, bottom) — (x, y, width, height) is missing y, width, height; (left, right, top, bottom) is missing right, top, bottom.
Looked for → y: y, top, top_left_y | width: width | height: height | right: IA_RIGHT, right, end_x | top: IA_TOP, top, start_y, top_left_y | bottom: IA_BOTTOM, bottom, end_y
Fields that did resolve: text='word', x='start_x', left='start_x'
Columns present in the words/IA table (4): subject, para, word, start_x
Matching ignores case and separators (IA_LEFT == ia_left == 'Ia Left') and takes the first candidate that matches; failing that, a vendor prefix or suffix on a known name (AOI_LEFT, LEFT_px) is tried next, accepted only when exactly one column qualifies.
To override auto-detection pass the full mapping, e.g. word_schema={'trial': '<column>', 'word_id': '<column>', 'x': 'start_x', 'y': '<column>', 'width': '<column>', 'height': '<column>'} — api.propose_schema(df, 'words') returns what was detected.
Pass the complete mapping the message's last line suggests — an explicit schema
replaces auto-detection wholesale — or start from
api.propose_schema(table, "words") and fill the gaps.
Rendering¶
fig = sps.plot_scanpath(words, fixations, pid, tid, canvas_size=(2560, 1440))
anim = sps.animate_scanpath(words, fixations, pid, tid, playback_speed=4.0)
pair = sps.compare_scanpaths(words, fixations, (pid, tid), (pid_b, tid_b))
sps.save_figure(fig, "out.html") # interactive, no browser needed
sps.save_figure(fig, "out.png") # .png/.svg/.pdf via Kaleido → needs Chrome
sps.save_figure_layers(fig, "layers/", fmt="svg") # one file per layer
HTML never needs Chrome. PNG/SVG/PDF go through Kaleido, which drives a
Chrome/Chromium binary: run plotly_get_chrome -y once, or fall back to HTML.
Figure options¶
Every figure keyword, its default, its render flag and the builders that
accept it: the figure options table, or
api.figure_options(kind) at runtime. The option values below are the ones
neither reference spells out.
color_by is a fixation column name ("duration_ms", "pass_index", a
pupil size you kept with load_scanpath_data(keep_columns=[…]), …),
the sentinel "(uniform)" for one flat color, or "line" to color each
fixation by the text line it lands on (the lines are inferred from word-box
geometry); a name the frame doesn't have raises a ValueError naming the
closest columns (see Errors). color_by_line=True
is the same as color_by="line", and on a single-trial figure it overrides any
other color_by. A comparison figure, and a co-animation (fixations_b=),
colors as the app's Compare mode does: each scanpath keeps its own color on
its marker outlines, while a numeric color_by fills both readings' markers on
one shared scale and a categorical one (or "line") on one shared
category→color mapping, with a legend entry per category.
fixation_flags marks or drops suspicious fixations (display only — reading
measures and exports are untouched). One entry per category, each with a mode of
"Off" / "Highlight" / "Discard":
flags = {
"short": {
"mode": "Highlight",
"threshold_ms": 80.0,
"symbol": "triangle-up-open",
"color": "#ff7f0e",
},
"long": {
"mode": "Off",
"threshold_ms": 800.0,
"symbol": "square-open",
"color": "#9467bd",
},
"oob": {"mode": "Discard", "symbol": "x", "color": "#d62728"}, # out of text
}
fig = sps.plot_scanpath(words, fixations, pid, tid, fixation_flags=flags)
saccade_color_mode is "Uniform", "Forward / regression" (the two-way fold)
or "By type" (forward / skip / refixation / return sweep / regression, each a
legended sub-trace, classified at render time);
saccade_class_colors={"regression": "#000", …} overrides individual class
colors. saccade_classes is the same split used as a filter rather than as
hue — saccade_classes=["regression"] draws a regressions-only figure (the
hidden classes lose their direction arrows too), and it composes with any
color mode; naming every class is a no-op. saccade_render_mode="Arc" draws
the linear-reading schematic.
heatmap_style is "Word boxes" or "Interpolated";
heatmap_metric="counts" weights by fixation count instead of dwell time;
heatmap_norm="Log" compresses heavy-tailed dwell times. Interpolated blurs the
fixations with a Gaussian of σ heatmap_sigma_px px — None (the default) is
2% of the data's larger span, at least 8 px.
fixation_color_range and heatmap_range are (min, max) pairs in the
metric's own units — for the word-box heatmap, dwell time (ms) or fixations per
word. Left at None each trial is scaled to its own values (the word-box
heatmap from 0), and a comparison shares one scale across A and B. Pass a range
to put every trial on the same scale. The "Interpolated" style scales its
density to its own peak and ignores heatmap_range.
highlight_column is a boolean words column (OneStop's critical span by
default); the default is skipped when absent, a column you name must exist.
fit_to_monitor=True frames the whole canvas_size; False crops to the data.
show_coordinate_grid=True overlays zero-anchored monitor-pixel coordinates;
coordinate_grid_spacing=None selects a readable 1/2/5×10ⁿ interval, while a
positive number pins the major interval in pixels. background_image places a
stimulus screenshot under the scanpath at data coordinates.
palette= is a shorthand that sets a whole group of colors at once —
"default" (colorblind-safe), "print" (grayscale) or "high-contrast"; the
app's own palette names work too (constants.PALETTES). Anything you pass explicitly still wins over it, and an
unknown name raises rather than silently falling back.
Headless defaults
plot_scanpath draws the app's default Scanpath design: fixations,
saccades and the text. Turn on show_words, show_heatmap or show_order
for word boxes, the heatmap or fixation numbers.
The same thing from the shell¶
scanpath-studio render --sample --list-trials
scanpath-studio render --sample -o scanpath.html
scanpath-studio render --words 'ia/*.csv' --fixations 'fix/*.csv' \
-p l37_1129 -t l37_1129_2_1_1_Ele_r0 --canvas 2560x1440 -o figure.png
scanpath-studio render --sample --animate --playback-speed 4 -o replay.html
Without -p / -t, render draws the first available trial instead of
raising. Every flag is in the CLI reference.
Errors and what they mean¶
| Message starts with | Cause | Fix |
|---|---|---|
Words/IA schema problems: / Fixations schema problems: |
A canonical field could not be inferred (or is missing from the schema you passed). | Read the bullets — they name the field, its schema key and the candidates tried. Pass word_schema= / fix_schema= built from api.propose_schema. |
Words/IA schema maps N column names the … table doesn't have |
A schema you passed names a column that isn't in the table. | The message lists each bad key and the closest real column names. |
fix_index_range=(a, b) selects no fixations |
The window is outside the trial. | The message gives the trial's fixation count and index range. |
words must be the normalized pandas DataFrame |
A path/string was passed where a frame belongs. | Run it through load_scanpath_data first. |
words frame is not normalized: |
A raw table (or a renamed frame) reached a plotting function. | Same — the frames the loader returns are the only accepted input. |
Ambiguous selection: N trials match |
participant / trial left out with several trials loaded. |
Pass both; list_trials shows what exists. |
No trial matches participant=… |
Unknown id. | The message lists available ids and the closest spellings. |
plot_scanpath() got an unexpected keyword argument |
Misspelled or unsupported option. | The message suggests the nearest names; api.figure_options() is the full list. |
color_by='…' (--color-by on the CLI) names no column (or highlight_column=, words) |
The option's value is a column the data doesn't have. | The message names the closest columns and lists them all; color_by also takes '(uniform)' and 'line'. |
Options not supported by the animation: |
A static-only option (heatmap, arcs, saccade types) passed to animate_scanpath. |
Drop it, or render the static figure. |
Static .png export failed: |
Kaleido has no Chrome. | plotly_get_chrome -y, or save .html. |
Fixations … have no usable coordinates |
AOI-sequence fixations with no matching word boxes. | Supply the words table whose word_ids match. |
GIF / MP4 of a replay¶
There is no api.py entry point for animated GIF or MP4. Use
animation_export.export_animation, which returns bytes and is keyword-only
(Kaleido + Chrome; ffmpeg rides along with imageio-ffmpeg). The clip lasts
what the replay does — the reading time over playback_speed — unless you pass
frame_duration_ms:
from pathlib import Path
from scanpath_studio.animation_export import export_animation
anim = sps.animate_scanpath(words, fixations, pid, tid, playback_speed=4.0)
clip = export_animation(anim, fmt="mp4") # or "gif"; lasts reading time / 4
Path("replay.mp4").write_bytes(clip)
Ground rules¶
- Errors name the alternatives. An unknown trial, an ambiguous selection or a misspelled option raises with the valid values listed; read the message rather than guessing again.
- Word boxes come from the data. They are never computed — only the fixation → word assignment is, and only when the data maps no word/IA id (box containment, else unassigned).