Skip to content

Python API

The public API follows one pipeline:

load data → list trials → plot → save
import scanpath_studio as sps

words, fixations = sps.load_scanpath_data("ia.csv", "fixations.csv")
fig = sps.plot_scanpath(words, fixations, participant="p1", trial="t3")
sps.save_figure(fig, "scanpath.html")

All functions below are importable from scanpath_studio. The figure options table lists every figure keyword.

Columns keep the names your files give them. The demo's word table calls its word id IA_ID and its total fixation duration IA_DWELL_TIME, and every function and option takes those names. A column Scanpath Studio made keeps its internal name. load_scanpath_data(..., names="canonical") gives the internal names, which are the same for every dataset. On the bundled demo, the first steps print this (run when the docs are built):

import scanpath_studio as sps

words, fixations = sps.load_sample_data()
print(sps.list_trials(words, fixations).head(3))

columns = ["IA_ID", "IA_LABEL", "IA_FIRST_FIXATION_DURATION", "IA_DWELL_TIME"]
print(words[columns].head(3))
  participant_id        unique_trial_id
0       l37_1129  l37_1129_2_1_1_Ele_r0
1       l37_1129  l37_1129_2_1_2_Ele_r0
2       l37_1129  l37_1129_2_1_3_Adv_r0
   IA_ID  IA_LABEL  IA_FIRST_FIXATION_DURATION  IA_DWELL_TIME
0      0    Robert                        32.0            387
1      1  Myslajek                       170.0            785
2      2     stops                       290.0           1330

Load

scanpath_studio.api.load_scanpath_data

load_scanpath_data(words: TablesLike | None = None, fixations: TablesLike | None = None, *, word_schema: dict | None = None, fix_schema: dict | None = None, trial_parts_manifest: dict | None = None, image_root: str | Path | None = None, image_pattern: str = '{text_id}.png', keep_columns: Iterable[str] | None = None, names: str = NAMES_SOURCE) -> ScanpathData

Load and normalize a words table and/or a fixations table.

The columns keep the names your files give them: CURRENT_FIX_DURATION, not duration_ms. A column Scanpath Studio built, converted, computed or changed keeps its internal name, and data.column_names (a ScanpathData) records what every column was called. Every API function takes these frames, and every column option (color_by=, hover fields …) takes either name. names="canonical" returns the internal names instead — the same for every dataset, for code that works across them.

words / fixations may be DataFrames, paths to .csv / .tsv / .txt / .tab / .parquet / .feather / .xlsx / .xls files (or a .zip of them), glob patterns, or lists of paths — multi-file datasets (one file per participant and/or text) are concatenated, with each file's stem kept in a source_file column. Column schemas are auto-detected (EyeLink, Gazepoint, Tobii, SMI, Pupil Labs, and snake_case names); pass word_schema / fix_schema mappings (field → column name; see propose_schema) to override detection.

trial_parts_manifest accepts a nested parent-trial/parts definition for datasets whose source tables identify screens through arbitrary selector columns; explicit screen_id / screen_index columns can instead be mapped directly in each schema. Either table may be omitted for datasets that ship only one report: the missing side comes back as an empty canonical frame and the plots simply skip that layer. Words without a participant column (stimulus-level AOIs) are copied onto every trial in the fixations — each trial matched by its trial id, else the trial id it had before a repeat's _r2 suffix, else its text_id (trial ids that embed the participant), with a data.StimulusJoinWarning (a UserWarning) when some trials match none — and fixations without x/y but with a word/AOI ID are placed at word-box centers. Columns named in data.INTERNAL_COLUMNS are the pipeline's bookkeeping (data.drop_internal_columns removes them).

Normalization keeps the mapped fields and the recognized optional ones (eye, EyeLink's interest-area measures, linguistic features …) and drops the rest. keep_columns names further columns of your own to carry through under their own names — a pupil size, a detection confidence — from whichever table has them, so a figure can color, hover or plot by them (the app's Extra fields to keep; render --keep-columns on the command line).

Returns the normalized (words, fixations) frames the plotting functions expect. Raises ValueError if a required field can't be found — the message names the canonical field, the column names auto-detection looked for, and the columns the table actually has — and data.StimulusJoinError (a ValueError) when a stimulus-level words table shares neither a trial id nor a text_id with any trial (or, multipart, with every screen a trial has fixations on).

scanpath_studio.api.ScanpathData

Bases: tuple

What load_scanpath_data returns: the (words, fixations) frames — so words, fixations = load_scanpath_data(…) unpacks it — plus column_names, the dataset's own name for every canonical column, per table ({"words": ColumnNames, "fixations": ColumnNames}).

With names="source" (the default) the frames' columns are the dataset's own names and each frame carries its map, so any API function takes it — sliced, filtered or merged — and the names it writes are yours. With names="canonical" they are the internal names, the same for every dataset.

scanpath_studio.api.load_sample_data

load_sample_data(*, names: str = NAMES_SOURCE) -> ScanpathData

Return the bundled OneStop demo, normalized and ready to plot: two participants, twelve trials each, every one of them with fixations. Under the demo's own column names; names="canonical" for the internal ones (see load_scanpath_data).

The frames carry the demo's recorded screen (OneStop's 2560×1440), so plot_scanpath draws them on it without a canvas_size, as scanpath-studio render --sample does.

scanpath_studio.api.load_raw_gaze

load_raw_gaze(table: TablesLike, *, raw_gaze_schema: dict | None = None, names: str = NAMES_SOURCE) -> DataFrame

Load and normalize a raw (sample-level) gaze table for raw_gaze=.

The third table plot_scanpath can draw, under the fixations: one row per eye-tracker sample, with a participant, a trial, x / y and usually a timestamp. table is a DataFrame, path, glob or list of paths, like load_scanpath_data's, and the columns are auto-detected the same way; pass raw_gaze_schema (field → column, see api.propose_schema(table, "raw_gaze")) to override the detection. plot_scanpath keeps only the plotted trial's (and screen's) samples, so one table can serve a whole corpus::

raw_gaze = sps.load_raw_gaze("gaze_samples.csv")
fig = sps.plot_scanpath(words, fixations, "p1", "t3", raw_gaze=raw_gaze)

Under the table's own column names, like load_scanpath_data's; names="canonical" for the internal ones.

scanpath_studio.api.load_sample_raw_gaze

load_sample_raw_gaze(*, names: str = NAMES_SOURCE) -> DataFrame

The bundled demo's raw gaze, normalized — what the app overlays on it.

OneStop ships no sample-level gaze, so this is synthesized from one of the demo's real trials and covers that trial alone.

scanpath_studio.api.load_participant_metadata

load_participant_metadata(table: TablesLike, *, id_column: str | None = None, participants: DataFrame | list | None = None)

Load a participant-level metadata table.

table is a DataFrame or a path/glob to a CSV/TSV/Parquet/Excel file with one row per participant: an id column plus anything known about them (native_language, age, a comprehension score). id_column defaults to the first recognized spelling (participant_id, subject, RECORDING_SESSION_LABEL, …).

Pass participants — a normalized frame or a list of ids — to have the join validated against the data you actually loaded; the returned object's .report then names the participants missing from either side.

Returns a ParticipantMetadata: the cleaned frame, a field registry (name, label, grain, dtype, missingness), and the join report. Nothing is broadcast onto the words/fixations frames — use scanpath_studio.metadata.project to attach chosen columns to a per-trial frame, or .values_for(pid) for one participant.

words, fixations = load_sample_data() meta = load_participant_metadata( ... "readers.csv", participants=fixations ... ) # doctest: +SKIP meta.names # doctest: +SKIP ('native_language', 'age')

scanpath_studio.api.load_trial_metadata

load_trial_metadata(table: TablesLike, *, id_column: str | None = None, participant_column: str | None = None, trials: DataFrame | None = None)

Load a trial-level metadata table.

The sibling of load_participant_metadata, one grain down: table has one row per trial — a trial-id column plus anything known about that trial (a list name, a condition, a per-trial comprehension score).

The key is yours to state, and it changes what the table means. Keyed by trial id alone, a row describes a text, and every trial of it inherits that row; pass participant_column to key by participant and trial, so a row describes one trial. Nothing in a file says which world a corpus is in, so this is never inferred — unlike id_column, which defaults to the first recognized spelling (trial_id, item_id, TRIAL_INDEX, …).

Pass trials — a normalized fixations/words frame, or any frame with participant_id + trial_id — to have the join validated against the data you actually loaded; the returned .report then names the trials missing from either side.

Returns a TrialMetadata: the cleaned frame, a field registry, and the join report. As with the participant table, nothing is broadcast onto the words/fixations frames.

words, fixations = load_sample_data() meta = load_trial_metadata( ... "readings.csv", trials=fixations ... ) # doctest: +SKIP meta.names # doctest: +SKIP ('list_name', 'comprehension_score')

scanpath_studio.api.load_text_metadata

load_text_metadata(table: TablesLike, *, id_column: str | list[str] | None = None, texts: DataFrame | list | None = None)

Load a text-level metadata table — the third grain.

table has one row per text — a text-id column plus anything known about that text (genre, difficulty, a stimulus-level comprehension score). Flat grain, like load_participant_metadata: never keyed by participant, since a text is a stimulus rather than something one participant owns. id_column defaults to the first recognized spelling (text_id, paragraph_id, stimulus_id, …) and may be several columns to build a composite id, the same way the uploaded data's own Text ID mapping does.

Pass texts — a normalized fixations/words frame, or any iterable of text ids — to have the join validated against the data you actually loaded; the returned .report then names the texts missing from either side.

Returns a TextMetadata: the cleaned frame, a field registry, and the join report. As with the other two grains, nothing is broadcast onto the words/fixations frames.

words, fixations = load_sample_data() meta = load_text_metadata( ... "texts.csv", texts=words ... ) # doctest: +SKIP meta.names # doctest: +SKIP ('genre', 'difficulty')

scanpath_studio.api.propose_schema

propose_schema(table: TablesLike, kind: str = 'words') -> dict

Auto-detected column mapping for a raw (un-normalized) table.

kind is "words", "fixations" or "raw_gaze". Returns {canonical field: source column or None} — the same mapping load_scanpath_data infers internally, so it's the place to start when detection got a field wrong or couldn't find one: edit the dict and pass it back as word_schema= / fix_schema=::

from scanpath_studio import api

schema = api.propose_schema("ia.csv", "words")
schema["trial"] = "TRIAL_LABEL"
words, fixations = api.load_scanpath_data("ia.csv", "fix.csv",
                                          word_schema=schema)

table is a DataFrame, path, glob or list of paths, like the loader's.

scanpath_studio.api.build_authored_scanpath

build_authored_scanpath(text: str, events: DataFrame | None = None, **layout_options) -> tuple[DataFrame, DataFrame]

Build normalized word/fixation frames from hand-authored reading events.

When events is omitted, one centered fixation per laid-out word is used. layout_options are forwarded to authoring.layout_text.

scanpath_studio.api.load_authored_scanpath

load_authored_scanpath(source: str | Path) -> tuple[DataFrame, DataFrame]

Load an authoring file — the JSON the app's Download authoring file saves — or its text, as normalized word/fixation frames.

scanpath_studio.datasets.load_potec

load_potec(root, *, readers: Iterable | None = None, texts: Iterable[str] | None = None, download: bool = False, names: str = 'source') -> tuple[DataFrame, DataFrame]

Load PoTeC as normalized (words, fixations) frames, ready to plot.

Under PoTeC's own column names; names="canonical" for the internal ones (see api.load_scanpath_data).

root is a clone of the PoTeC repo (with the eye-tracking data downloaded) or any folder; with download=True the needed files are fetched into it on first use (~10 MB). Narrow the load with readers (e.g. [0, 1]) and/or texts (e.g. ["b0", "p3"]) — the full corpus is 75 readers × 12 texts = 900 trials.

Participants are PoTeC reader ids (as strings); a trial is one reader's reading of one text, <reader>_<text> ("0_b0"), and text_id is the text (b0–b5 biology, p0–p5 physics)::

import scanpath_studio as sps

words, fixations = sps.load_potec("data/PoTeC", readers=[0], texts=["b0"])
fig = sps.plot_scanpath(
    words, fixations, "0", "0_b0", canvas_size=(1680, 1050)
)

The PoTeC monitor was 1680×1050 (DELL P2210, 60 Hz); pass that as canvas_size to plot_scanpath for true-to-scale rendering.

scanpath_studio.datasets.load_onestop

load_onestop(root, *, regime: str = 'ordinary', parts: Iterable[str] | None = None, variant: str = 'public', download: bool = False, names: str = 'source') -> tuple[DataFrame, DataFrame]

Load OneStop as normalized (words, fixations) frames, ready to plot.

Under OneStop's own column names; names="canonical" for the internal ones (see api.load_scanpath_data).

root is a folder holding (or to download into, public variant only) the OneStop reports. Narrow the load with regime (ordinary / information_seeking / repeated / information_seeking_repeated), parts (any subset of Title / Question_Preview / Paragraph / Questions / Answers / QA / Feedback — default Paragraph), and variant (public OSF release or lacclab local export). The public OSF reports are large; pass download=True to fetch the chosen regime + parts into root on first use::

import scanpath_studio as sps

words, fixations = sps.load_onestop(
    "data/OneStop", regime="ordinary", parts=["Paragraph"], download=True
)
pid, tid = sps.list_trials(words, fixations).iloc[0]  # one reading
fig = sps.plot_scanpath(words, fixations, pid, tid, canvas_size=(2560, 1440))

OneStop's presentation monitor was 2560×1440 px (Dell U2715H; Berzak et al. 2025, Methods → Apparatus); pass that as canvas_size to plot_scanpath for true-to-scale rendering. The reports already match the bundled demo's schema, so this reuses the generic auto-detect → normalize path (no OneStop-specific column mapping).

Inspect

scanpath_studio.api.list_trials

list_trials(words: DataFrame | None = None, fixations: DataFrame | None = None, *, raw_gaze: DataFrame | None = None) -> DataFrame

One row per plottable trial: its participant id and trial id.

Trials present in both frames when both are loaded; for single-report datasets (words-only or fixations-only), trials from whichever frame has data. raw_gaze (a frame from load_raw_gaze) adds the trials that only its samples cover — every trial, for a dataset recorded as raw gaze alone (pass None for words and fixations then). The id columns take the names the frames carry.

scanpath_studio.api.list_parts

list_parts(words: DataFrame | None, fixations: DataFrame | None, participant: str | None = None, trial: str | None = None, *, raw_gaze: DataFrame | None = None) -> DataFrame

Ordered screens in multipart data, optionally narrowed to one parent.

Single-screen data returns an empty table. A trial recorded as raw gaze alone takes its screens from raw_gaze (its screen_id), decided per trial — so a samples-only trial keeps its screens in a dataset whose other trials have fixations.

scanpath_studio.api.check_data_health

check_data_health(words: DataFrame | None = None, fixations: DataFrame | None = None, raw_gaze: DataFrame | None = None) -> DataFrame

Values that loaded as numbers but cannot be right — the Data Management page's Data checks.

Checks the normalized tables (from load_scanpath_data / load_raw_gaze) for fixations lasting 0 ms or less or with an infinite duration or onset, fixations and raw-gaze samples whose position is missing or infinite, word boxes with no area or no finite position, and per-screen screen sizes that are not finite and positive. One row per check that found anything: table, check, problem, the columns it read (in the names the frames carry), rows of of_rows, the trials they fall in, a breakdown by kind, severity ("note" for raw-gaze gaps, which blinks and track loss make ordinary), what_happens to those rows in the app, and a few examples. An empty frame means every check passed. Nothing is changed or dropped::

words, fixations = sps.load_scanpath_data("ia.csv", "fixations.csv")
print(sps.check_data_health(words, fixations))

Plot

scanpath_studio.api.plot_scanpath

plot_scanpath(words: DataFrame | None = None, fixations: DataFrame | None = None, participant: str | None = None, trial: str | None = None, *, screen: str | None = None, canvas_size: tuple[int, int] | None = None, base_font_size: int = 16, font_family: str = FONT_FAMILY, raw_gaze: DataFrame | None = None, drift_correction: str | None = None, drift_connectors: bool = False, fix_index_range: tuple[int, int] | None = None, illustration: bool = False, illustration_label: str = 'auto', title: str = '', caption: str = '', column_names: dict | None = None, **figure_overrides) -> Figure

Build one trial's scanpath figure (by default the app's Scanpath design).

words / fixations are normalized frames from load_scanpath_data. participant / trial may be omitted when the frames hold exactly one trial. canvas_size is the monitor size in px; by default it is estimated from the data extents — pass the real monitor resolution (e.g. (2560, 1440) for OneStop) to keep coordinates true to scale. For a multipart trial, screen selects one child screen; omitting it selects the first recorded screen and never concatenates coordinate spaces. raw_gaze is a frame from load_raw_gaze, filtered to the selected trial and drawn as recorded. It can be the only table: for a dataset recorded as raw gaze alone pass None for words and fixations (plot_scanpath(raw_gaze=samples, trial=…)) — the trial is looked up in the samples, the canvas is estimated from their extent, and the figure is the samples alone. Nothing is derived from them: no fixations are detected, so the fixation, saccade and heatmap layers stay empty.

drift_correction / drift_connectors are experimental: without SCANPATH_EXPERIMENTAL=1 any drift_correction other than None / "off" raises ValueError.

fix_index_range=(start, end) draws only fixations start through end (1-based, both inclusive) of the trial — the headless form of the app's fixation-index window.

title / caption stamp a title/caption band onto the figure without shrinking the plot area, like the app's Title & labels — literal text here, not the app's {trial_id}-style pattern, since the caller already knows which trial this is.

illustration=True applies the Illustration preset (snapped fixations, arced saccades, uniform colors, no heatmap or word boxes); keywords you pass still win. illustration_label is "auto" (label the figure when it no longer shows the data as recorded), "show" or "hide". palette= ("default", "print" or "high-contrast", or the app's names) sets a group of colors at once; a color you pass explicitly wins.

Remaining keywords override the app's defaults and are forwarded to plots.make_scanpath_figure (e.g. show_heatmap=True, color_by="pass_index", x_field="order_in_trial"); an unknown keyword raises a TypeError naming the closest valid options, and figure_options lists them all with their defaults (choices=True adds the values each enumerated option takes; a value is matched ignoring case, spaces, - and _, and any other value raises a ValueError). A color_by / highlight_column naming a column the trial's table doesn't have raises a ValueError naming the closest ones, rather than drawing without it.

Frames under the dataset's own column names (what load_scanpath_data returns by default) are read through the map they carry, and an option naming a column takes either name. column_names is that map for frames loaded with names="canonical" (data.column_names): the options then take the dataset's names too, and the figure's text uses them.

scanpath_studio.api.animate_scanpath

animate_scanpath(words: DataFrame | None = None, fixations: DataFrame | None = None, participant: str | None = None, trial: str | None = None, *, screen: str | None = None, screen_b: str | None = None, canvas_size: tuple[int, int] | None = None, base_font_size: int = 16, font_family: str = FONT_FAMILY, playback_speed: float = 1.0, autoplay: bool = True, fix_index_range: tuple[int, int] | None = None, fix_index_range_b: tuple[int, int] | None = None, illustration_label: str = 'auto', title: str = '', caption: str = '', column_names: dict | None = None, trial_b: tuple[str, str] | None = None, dataset_b: str | None = None, setup: SetupSnapshot | None = None, setup_b: SetupSnapshot | None = None, **animation_overrides) -> Figure

Build the animated scanpath replay for one trial.

Same trial selection, canvas and column-name semantics as plot_scanpath (column_names included), and screen selection for multipart trials. The replay takes the reading time divided by playback_speed: save it as interactive HTML with save_figure, whose page keeps that clock itself, or rasterize it to GIF/MP4 with animation_export.export_animation, which lasts as long. (fig.show() plays it on Plotly's own frame queue, which runs slow.) fix_index_range=(start, end) replays only that window of the trial's fixations (1-based, inclusive), like plot_scanpath.

With autoplay (default True) the saved interactive HTML auto-starts the replay on load at playback_speed — save_figure honors the marker the builder stamps on the figure. Pass autoplay=False to save a figure that opens paused (press ▶ Play to run it). Autoplay only affects the interactive HTML; a GIF/MP4 always plays from its first frame.

When playback_speed is not 1, the automatic Illustration label says the replay timing was changed. illustration_label accepts "auto", "show", or "hide" like plot_scanpath.

In a co-animation fix_index_range windows A only (the app's rule — A's slider never cuts B), fix_index_range_b windows B, and fixation_flags_b gives B flags of its own (None: A's fixation_flags, or the fixation_flags of style_b when it names some).

style_a / style_b style the two scanpaths of a co-animation as they style compare_scanpaths' — the same keys (fix_color, marker_size_range, opacity, hollow, saccade_color, saccade_style, saccade_width), resolved the same way, so the replay and the static comparison draw each trial alike. The replay has no saccade-class filter, so a style naming saccade_classes raises ValueError. A lone replay ignores both.

trial_b=(participant, trial) co-animates a second trial on the same clock, like the app's Animate + Compare. It is looked up in words_b / fixations_b when given, else in words / fixations — the way compare_scanpaths takes it. Without trial_b, words_b / fixations_b must hold one trial; B frames holding several raise ValueError rather than drawing them all. A multipart B is drawn at screen_b — looked up in B's own trial — or at its first recorded screen without it, as A is with screen.

Two datasets. Both readings are drawn in A's coordinates, so a co-animation is an overlay, and a trial from another dataset has to share A's screen. Name that dataset with dataset_b (or give its setup_b) and the pair is checked the way compare_scanpaths checks an overlay: two different canvases raise IncomparableScreensError, a ValueError, rather than draw. setup / setup_b are experimental_setup.SetupSnapshot values; a side without one is read off its data — the extent of that one trial, which rarely spans the whole screen, so state both when you know them — and canvas_size covers A when you only have a resolution. dataset_b also prefixes B's participant ids with the dataset's name, as compare_scanpaths does, so a hover says whose participant it is. words_b / fixations_b passed without either are taken to be from A's dataset, as render passes them for --compare-with alone, and are not checked: two readings of one corpus can span different extents, and inferring a canvas from each would refuse pairs that shared a screen.

The animation builder accepts a subset of the static figure's options (show_words, show_word_labels, show_saccades, show_order, styling, and second-scanpath overlays) — see figure_options("animation"); an unsupported key raises a ValueError naming the valid ones. The shared options default to the same values as plot_scanpath (CANONICAL_FIGURE_DEFAULTS), so the replay matches the static figure. palette= works here too; the colors it implies that the animation doesn't support are dropped rather than raising, since the caller named a look, not those individual keys.

title / caption — same as plot_scanpath.

The replay is made of fixations, so a trial without any — one recorded as raw gaze alone, or a words-only one — raises ValueError rather than returning an empty replay, and raw_gaze= is refused: the replay draws no raw-gaze layer, and nothing detects fixations from samples. Draw samples with plot_scanpath(raw_gaze=…).

scanpath_studio.api.compare_scanpaths

compare_scanpaths(words: DataFrame, fixations: DataFrame, trial_a: tuple[str, str], trial_b: tuple[str, str], *, screen: str | None = None, screen_b: str | None = None, words_b: DataFrame | None = None, fixations_b: DataFrame | None = None, dataset_b: str = 'Dataset B', raw_gaze: DataFrame | None = None, raw_gaze_b: DataFrame | None = None, layout: str = 'overlay', compare_stimulus: str = 'both', setup: SetupSnapshot | None = None, setup_b: SetupSnapshot | None = None, canvas_size: tuple[int, int] | None = None, labels: tuple[str, str] | None = None, style_a: dict | None = None, style_b: dict | None = None, base_font_size: int = 16, font_family: str = FONT_FAMILY, fix_index_range: tuple[int, int] | None = None, fix_index_range_b: tuple[int, int] | None = None, drift_correction: str | None = None, title: str = '', caption: str = '', column_names: dict | None = None, **figure_overrides) -> Figure

Build a two-scanpath comparison figure.

The headless form of the app's Compare mode. trial_a / trial_b are (participant, trial) pairs; layout is "overlay", "side_by_side" ("side-by-side" also accepted) or "stacked".

Multipart trials. Each scanpath is one screen, never a whole multipart trial: every screen is its own coordinate space, so pooling them would draw saccades across page boundaries. screen picks A's screen and screen_b B's, independently — B's is looked up in B's own frames, so it may be a later page or another dataset's. Either one left out is that trial's first recorded screen, as in plot_scanpath; list_parts() lists them. A screen named for a single-screen trial, or one the trial does not have, raises ValueError.

Two datasets. Pass words_b / fixations_b to draw B from a different dataset. Two datasets can hold the same (participant_id, trial_id) and the builder slices by exactly that pair, so B's participant ids are namespaced with dataset_b inside the throwaway merged frames — without it one trial would silently render as two. The frames you pass in are never modified, and nothing in the returned figure's data depends on the namespace beyond the trace labels.

The overlay gate. Across datasets an overlay needs both canvases to be the same size; otherwise this raises ValueError (the app falls back to side by side). One dataset can hold screens of different sizes too, so a same-dataset pair is refused the same way when the two selected screens carry different canvases (canvas_width / canvas_height columns) or setup_b states another screen. Pass layout="side_by_side" or "stacked" to compare readings from different screens; each panel is then drawn to its own. Nothing is rescaled.

setup / setup_b are experimental_setup.SetupSnapshot values — what the gate reads. canvas_size covers A when you only have a resolution; omit both and the canvas is read off the data.

Stimulus images. background_image is A's page. A split layout draws B's panel over background_image_b (with background_image_size_b / background_image_origin_b) and over nothing without it — never A's, since sharing a dataset says nothing about sharing a page.

compare_stimulus picks whose word boxes and text an overlay draws — "both" (default), "a" or "b". Two datasets' AOIs coincide only when the text is identical. Split layouts ignore it; each panel owns its own stimulus.

Per-scanpath style. style_a / style_b restyle one scanpath: fix_color, marker_size_range, opacity, hollow, saccade_color, saccade_style, saccade_width, box_color — the outline of that reading's word boxes, its fix_color when left out — box_fill_color, their fill, word_box_fill_color when left out — and raw_gaze_color, that reading's raw-gaze samples, its fix_color when left out. These three are this figure's only: the co-animation draws one set of boxes, in word_box_color / word_box_fill_color, and no raw gaze, and ignores them. heatmap_colorscale gives that reading's word-box heatmap its own color scale (heatmap_colorscale when left out) on the range both share; when A's and B's differ, each gets its own color bar.

Filters, per scanpath. fixation_flags and saccade_classes filter both scanpaths, as they filter plot_scanpath's one; the same two keys in style_a / style_b give that scanpath its own, overriding them — e.g. style_b={"fixation_flags": {"short": {"mode": "Discard", "threshold_ms": 80}}, "saccade_classes": ["regression"]}. The app's Compare mode draws A under the plot controls' filters and B under its own. fix_index_range windows both scanpaths; fix_index_range_b gives B a window of its own (the app's B slider).

Raw gaze. raw_gaze is a frame from load_raw_gaze; each reading's samples are drawn under its scanpath, in that scanpath's color (raw_gaze_marker_size / raw_gaze_opacity style them). It serves both readings of a same-dataset comparison; across datasets it is A's, and raw_gaze_b is B's. Passing either turns the layer on; show_raw_gaze=False keeps it off.

Remaining keywords are forwarded to plots.make_comparison_figure (e.g. show_words=False, color_by="duration_ms"); an unknown one raises TypeError naming the closest valid options; figure_options("comparison") lists the accepted keywords. Column names follow plot_scanpath's rule: A's names (or column_names) name the options and the figure's text, and either dataset's frames may come under their own names.

scanpath_studio.api.render_parent_trial

render_parent_trial(words: DataFrame, fixations: DataFrame, participant: str | None = None, trial: str | None = None, *, animate: bool = False, transition_mode: str = 'instant', screens: Sequence[str] | None = None, **options) -> dict[str, Figure]

Render every screen of one logical trial without stitching coordinates.

screens renders only those screen ids (one id may be given as a string), in the trial's own order, as the app's Export → Screens does; an id the trial does not have raises ValueError. None renders them all. screen_index in each figure's meta stays the screen's place in the trial.

The ordered mapping is keyed by screen_id. Each value is the same figure returned by plot_scanpath or animate_scanpath; callers can save them into deterministic per-screen files. transition_mode is "instant" or "recorded". For animated output, each figure's layout.meta['transition_after_ms'] records the delay before the next screen (zero for instant mode, or the observed parent-clock gap). No visual saccade is ever drawn across the boundary.

scanpath_studio.api.plot_corpus_figure

plot_corpus_figure(data: DataFrame, *, kind: str, measure_label: str = 'Value', series_col: str = 'series', value_col: str = 'value', colors: tuple[str, ...] | None = None, canvas_width: int = 1000, base_font_size: int = 14, font_family: str = FONT_FAMILY) -> Figure

Headless corpus profile/distribution/difference plot with shared colors.

profile expects word_id plus value_col (and optional lo / hi); distribution expects value_col; difference expects word_id and diff. When series_col is present, it defines the overlaid profile/distribution series. A table missing a column its kind reads raises ValueError naming it and the columns present.

Reproduce a figure in code

The app's Share subtab shows the API or CLI code that rebuilds the figure currently on screen — paste it into a notebook or terminal to get the same figure. figure_code is the headless form of that block, and render --print-code prints it for an invocation you already have.

print(
    sps.figure_code(
        participant="l7_1090",
        trial="l7_1090_2_1_1_Ele_r0",
        show_heatmap=True,
        color_by="duration_ms",
    )
)
import scanpath_studio as sps

words, fixations = sps.load_sample_data()

fig = sps.plot_scanpath(
    words,
    fixations,
    participant='l7_1090',
    trial='l7_1090_2_1_1_Ele_r0',
    canvas_size=(2560, 1440),
    color_by='duration_ms',
    show_heatmap=True,
)

sps.save_figure(fig, 'scanpath.png')

scanpath_studio.api.figure_code

figure_code(*, kind: str = 'static', source: str = 'demo', source_options: dict | None = None, participant: str = '', trial: str = '', screen: str | None = None, compare: tuple[str, str] | None = None, compare_screen: str | None = None, compare_layout: str = 'overlay', compare_stimulus: str = 'both', compare_dataset: str = '', compare_canvas: tuple[int, int] | None = None, compare_labels: tuple[str, str] | None = None, canvas_size: tuple[int, int] | None = None, base_font_size: int = 16, font_family: str = FONT_FAMILY, title: str = '', caption: str = '', fix_index_range: tuple[int, int] | None = None, illustration_label: str = 'auto', drift_correction: str | None = None, drift_connectors: bool = False, playback_speed: float = 1.0, autoplay: bool = True, flavor: str = 'python', explicit: bool = False, output: str | None = None, **figure_overrides) -> str

The API or CLI code that reproduces a figure.

The headless twin of the app's 🔗 Share → Reproduce this figure in code block: give it the same arguments you would give plot_scanpath (kind="static"), animate_scanpath ("animation") or compare_scanpaths ("comparison") and it returns the snippet that rebuilds that figure, rather than the figure::

print(sps.figure_code(participant="l7_1090", trial="l7_1090_2_1_1_Ele_r0",
                      show_heatmap=True, flavor="cli"))

source names how the data is loaded — "demo", "synthetic", "files", "potec", "onestop", "author", or "unknown" for data a snippet can't name — with source_options carrying that loader's arguments ({"root": …}, {"words": [...], "fixations": [...]}, and so on). With show_raw_gaze=True the raw-gaze table is read too: the demo's own, or the path(s) given as source_options["raw_gaze"] (plus an optional "raw_gaze_schema") — load_raw_gaze in the Python form, --raw-gaze in the CLI one. source="raw_gaze" is a dataset recorded as raw gaze alone: the samples at source_options["raw_gaze"] are the data, and plot_scanpath is handed None for the words and fixations.

screen / compare_screen are A's and B's screens of a multipart trial (screen= / screen_b=, --screen / --compare-screen).

compare_dataset names the dataset scanpath B was loaded from when it is a second one. B's participant id belongs to that dataset rather than the one the snippet loads, so both forms then load B's own tables and name B in them — words_b= / fixations_b= / dataset_b=, and --compare-words / --compare-fixations beside --compare-with — from the placeholder paths B_WORDS / B_FIXATIONS, which you point at its files. compare_canvas is B's screen, (width, height), when you know it: written as setup_b= and --compare-canvas, which a co-animation across datasets needs.

compare_labels is the pair you would pass compare_scanpaths as labels= — the two trace labels, when they are not the composed defaults. Both forms carry them: labels= in the Python snippet, --label-a / --label-b in the CLI one.

With participant / trial left empty the snippet renders the first available trial, as render does. canvas_size defaults to the screen render assumes for the source (the demo's 2560×1440, PoTeC's 1680×1050, …), so both flavors draw the same figure; output defaults to scanpath.html for an animation — render --animate writes only HTML — and to a PNG otherwise.

Only the options that differ from figure_options are written, so the snippet stays readable; explicit=True emits every option at its current value. flavor is "python", "cli", or "both" (the two separated by a blank line). Anything neither form can reproduce — a raw-gaze table with no path to name, an uploaded stimulus image, B's rows from a second corpus — follows as # Note: comments, matching the ⚠️ captions the app shows and the Note: lines render --print-code writes to stderr. See code_snippet.ReproductionCode for the structured form.

scanpath_studio.api.figure_options

figure_options(kind: str = 'static', *, choices: bool = False) -> dict

Every figure keyword a builder accepts → the default it renders with.

With choices=True each name maps to {"default": …, "choices": …}, where choices is the tuple of values an enumerated option takes (heatmap_norm: ("Linear", "Log")) and None for a free one. An enumerated option takes any spelling of a choice — case, spaces, - and _ are ignored, so the CLI's "log" and "mark-border" work — and raises ValueError listing them for anything else.

kind="static" covers plot_scanpath, kind="animation" animate_scanpath (whose builder supports a subset), and kind="comparison" compare_scanpaths. The values are the defaults a call actually renders with (the app's Scanpath design) — so a scripted caller can diff its intended settings against what it would get::

{k: v for k, v in sps.figure_options().items() if k.startswith("show_")}

Figure options

Every keyword the figure builders take, with the default it renders with, the values it takes when there is a fixed set (any case, and - or _ for a space, so the CLI's "log" and "mark-border" work; anything else raises a ValueError), the render flag that sets it on the command line, and which builders accept it: plot is plot_scanpath, animate is animate_scanpath, compare is compare_scanpaths.

Option Default Values render flag Accepted by
anim_grid_step_ms None — --anim-grid-step-ms animate
anim_max_frames None — --anim-max-frames animate
background_color '#ffffff' — --background-color all three
background_image None — --stimulus-image all three
background_image_b None — --stimulus-image-b compare
background_image_opacity 1.0 — --stimulus-image-opacity all three
background_image_origin None — --stimulus-image-origin all three
background_image_origin_b None — --stimulus-image-origin-b compare
background_image_size None — --stimulus-image-size all three
background_image_size_b None — --stimulus-image-size-b compare
color_by '(uniform)' — --color-by all three
color_by_line False — --color-by-line all three
compare_stimulus 'both' 'both', 'a', 'b' --compare-stimulus animate
connector_y None — — plot, compare
coordinate_grid_spacing None — --coordinate-grid-spacing all three
critical_span_style 'Mark text' 'Mark text', 'Mark border', 'None' --critical-span-style plot, compare
duration_size_legend True — --no-duration-size-legend all three
fit_to_monitor True — --no-full-monitor all three
fixation_color '#0072B2' — --fixation-color all three
fixation_color_range None — --fixation-color-range all three
fixation_colorbar_orientation 'Vertical' 'Vertical', 'Horizontal' --fixation-colorbar-orientation all three
fixation_colorbar_tickangle 0 — --fixation-colorbar-tickangle all three
fixation_colorbar_tickfont_size 12 — --fixation-colorbar-tickfont-size all three
fixation_colorscale 'Blues' — --fixation-colorscale all three
fixation_flags None — --fixation-flag all three
fixation_flags_b None — --compare-fixation-flag animate
fixation_hover_fields ['order_in_trial', 'duration_ms', 'word_id'] — --fixation-hover-fields all three
fixation_opacity 0.7 — --fixation-opacity all three
fixation_snap_to_word False — --snap-fixations plot, compare
fixation_symbol 'circle' 'circle', 'square', 'diamond', 'triangle-up', 'cross', 'x', 'star', 'hexagon', 'heart' --fixation-symbol all three
fixations_b None — — animate
heatmap_colorbar_orientation 'Vertical' 'Vertical', 'Horizontal' --heatmap-colorbar-orientation all three
heatmap_colorbar_tickangle 0 — --heatmap-colorbar-tickangle all three
heatmap_colorbar_tickfont_size 12 — --heatmap-colorbar-tickfont-size all three
heatmap_colorscale 'Blues' — --heatmap-colorscale plot, compare
heatmap_metric 'duration_ms' — --heatmap-metric plot, compare
heatmap_norm 'Linear' 'Linear', 'Log' --heatmap-norm plot, compare
heatmap_range None — --heatmap-range plot, compare
heatmap_sigma_px None — --heatmap-sigma plot, compare
heatmap_style 'Word boxes' 'Word boxes', 'Interpolated' --heatmap-style plot, compare
highlight_column 'is_in_aspan' — --highlight-column all three
highlight_text_color '#D55E00' — --highlight-text-color all three
hollow_fixations False — --hollow-fixations all three
illustration_reasons None — — plot, compare
illustration_text '' — --illustration-text all three
label_a 'Scanpath A' — --label-a animate
label_b 'Scanpath B' — --label-b animate
legend_layout None — --legend all three
line_spacing 3.0 — --line-spacing all three
marker_duration_range (50, 600) — --marker-duration-range all three
marker_size_range (8, 24) — --marker-size-range all three
marker_size_scale 'sqrt' 'sqrt', 'linear', 'log', 'relative' --marker-size-scale all three
order_font_color '#111111' — --order-font-color all three
order_font_size 10 — --order-font-size all three
raw_gaze_color '#888888' — --raw-gaze-color all three
raw_gaze_marker_size 4.0 — --raw-gaze-marker-size all three
raw_gaze_opacity 0.6 — --raw-gaze-opacity all three
saccade_class_colors None — --saccade-type-color plot, compare
saccade_classes list (see figure_options()) — --saccade-classes plot, compare
saccade_color '#CC79A7' — --saccade-color all three
saccade_color_mode 'Uniform' 'Uniform', 'Forward / regression', 'By type' --saccade-color-by-type, --saccade-color-by-direction plot, compare
saccade_render_mode 'Straight' 'Straight', 'Arc' --saccade-arcs plot, compare
saccade_style 'solid' 'solid', 'dash', 'dot', 'dashdot' --saccade-style all three
saccade_type_legend True — --no-saccade-type-legend plot, compare
saccade_width 2.0 — --saccade-width all three
scale_text_to_boxes True — --no-scale-text-to-boxes all three
show_connectors False — — plot, compare
show_coordinate_grid False — --coordinate-grid all three
show_fixation_colorbar True — --no-fixation-colorbar all three
show_fixations True — --no-fixations plot, compare
show_heatmap False — --no-heatmap, --heatmap plot, compare
show_heatmap_colorbar True — --no-heatmap-colorbar all three
show_legend True — --no-compare-legend animate, compare
show_order False — --no-fixation-index, --fixation-index all three
show_raw_gaze False — --raw-gaze, --no-raw-gaze plot, compare
show_saccade_arrows False — --saccade-arrows all three
show_saccades True — --no-saccades all three
show_word_labels True — --no-text all three
show_words False — --no-word-boxes, --word-boxes all three
span_border_color '#000000' — --span-border-color plot, compare
style_a None — --style-a animate, compare
style_b None — --style-b animate, compare
text_color '#000000' — --text-color all three
word_box_color '#6c757d' — --word-box-color all three
word_box_fill_color '#646464' — --word-box-fill-color all three
word_box_fill_opacity 0.05 — --word-box-fill-opacity all three
word_box_line_opacity 1.0 — --word-box-line-opacity all three
word_heatmap_col None — --word-heatmap-col plot, compare
word_heatmap_title None — --word-heatmap-title plot, compare
word_hover_fields list (see figure_options()) — --word-hover-fields all three
word_hover_measure 'total_fixation_duration_ms' — --word-hover-measure all three
words_b None — — animate
x_field 'x' — --x-field plot, compare
y_field 'y' — --y-field plot, compare

Save

scanpath_studio.api.save_figure

save_figure(fig: Figure, path: str | Path, *, scale: float = 2, width: int | None = None, height: int | None = None, width_mm: float | None = None, width_in: float | None = None, dpi: int | None = None) -> Path

Save a figure by extension: .html (interactive, needs no browser) or .png/.svg/.pdf (static via Kaleido — needs Chrome, Chromium or Edge; run plotly_get_chrome -y once if none is installed). width / height set the image size in px (overriding the figure's own size); both ignored for .html. Returns the written path.

width_mm or width_in with dpi (default 300) sizes a PNG for print, as the app's Export → Current figure does: 180 mm at 600 dpi is a 4,252 px wide PNG, its height following the figure's aspect, with the dpi written into the file. They replace scale.

scanpath_studio.api.save_figure_layers

save_figure_layers(fig: Figure, directory: str | Path, *, fmt: str = 'svg', scale: int = 2, width: int | None = None, height: int | None = None) -> dict

Split a scanpath figure into its layers and save one file per layer.

Writes <directory>/<layer>.<fmt> for each visible layer (word boxes / fixations / saccades / heatmap / labels / stimulus image / frame) and returns {layer: Path}. Each layer is the full figure with only that layer's elements and a transparent background, at the same size and axis ranges — so the files register perfectly when stacked in Illustrator / Inkscape. fmt is any save_figure extension without the dot (svg / pdf are vector and best for editing; png / html also work). scale / width / height are forwarded to save_figure.

Recovery cache

scanpath_studio.api.cache_status

cache_status() -> dict

Describe the on-device recovery cache a local app run keeps.

The app stores completed uploaded datasets, column mappings, view settings, saved designs, metadata tables and annotations under the user's cache directory so a refresh or restart resumes where it left off — on localhost/desktop only, never on a hosted deployment. This reports that store without launching the app: enabled, directory, datasets (name + per-frame row counts), rows, annotations, designs, metadata, settings, bytes, saved_at, plus exists / readable for a missing or unreadable manifest, damaged (name + reason) for a stored dataset whose entry or files are broken — the app restores the others and keeps that one in the cache rather than dropping it — and damaged_metadata, the reason the stored metadata tables would not restore ("" when they would). Delete it with clear_cache; the same information is in the app's 🗂️ Data Management → Saved on this computer section and in scanpath-studio cache.

scanpath_studio.api.clear_cache

clear_cache() -> dict

Delete the on-device recovery cache and return its status afterwards.

Removes only the files this app wrote (manifest.json and the dataset Parquet files); anything else in the folder is left alone. A running local app writes its session back out at the end of its next change — start it with scanpath-studio run --no-persist or SCANPATH_STUDIO_PERSIST=0 to stop that.

Version and updates

scanpath_studio.api.version_info

version_info() -> BuildInfo

Which build of Scanpath Studio this is — no network access.

version is what scanpath_studio.__version__ holds: the release itself ("0.35.0"), or between releases a PEP 440 version that sorts after it — "0.35.0.post3+g8f18219" is three commits after v0.35.0, at commit 8f18219, and it ends .dirty with uncommitted changes. release is the release it descends from (scanpath_studio.__release__), distance the commits since (None when unknown), commit, dirty, and source — how it was worked out: "checkout" (git describe), "stamp" (a desktop bundle's build stamp), "vcs" (a pip install git+…) or "release". describe() says it in a sentence. The same is in Help → About and scanpath-studio version.

scanpath_studio.api.check_for_updates

check_for_updates(timeout: float = 5.0) -> UpdateCheck

Ask GitHub whether a newer release than this build is out.

The one call here that uses the network, and only when made: it reads the latest release from api.github.com (drafts and pre-releases excluded) and compares it with version_info. It never raises. status is "up_to_date", "update_available", "ahead" (a development build past the latest release) or "error" (offline, no answer within timeout seconds, rate-limited, …), and message says it in a sentence. With an update available, command is the shell command that updates this install (pip install -U scanpath-studio, uv tool upgrade scanpath-studio, git pull, …), latest.url the release notes, and in the desktop app download the archive for this computer. The same check is Help → About → Check for updates and scanpath-studio version --check.

For a batch loop, see Automation. GIF and MP4 export uses scanpath_studio.animation_export.export_animation and requires Kaleido plus Chrome, Chromium or Edge.