Skip to content

Visualisation API

Visualisation helpers turn saved run histories into satisfaction plots and standalone legends for notebooks, reports, and paper figures.

Episode Metrics

Aggregate per-agent episode rows into one algorithm-level row per episode.

Build one algorithm-level completed-episode metric series for one run.

Plot an algorithm-level completed-episode metric against episode end step.

Write titleless completed-episode graph variants for a suite.

Satisfaction Plots

One panel, one curve PER AGENT, showing EWMA-smoothed success.

Input
  • history_by_agent (single run) OR
  • runs = [history_by_agent_seed0, history_by_agent_seed1, ...]

success per episode is inferred from EpisodeLog[reward_key] > success_threshold. For your wrapper, cum_reward is 0/1, so threshold=0 is correct.

Tips
  • 3 runs => band="ci95" is meaningful.
  • x_axis="episode", align="truncate" gives clean episode-indexed CI bands.
  • alpha: smaller => smoother (slower), larger => more responsive. Rough intuition: effective window ≈ 1/alpha episodes.

If save_path is set, the base file is saved without an inline legend and save_legend_variant=True also writes a sibling "*_legend.png" file.

One panel, one curve PER ALGORITHM VARIANT. Base and CER variants are both shown when present.

reduce_agents
  • "mean": average agent EWMA satisfaction within each run.
  • "max": use the best agent EWMA satisfaction at each episode.

agent: when set, plot only that agent's EWMA satisfaction.

include_handcrafted: when true, add the best handcrafted-reward result as a single endpoint marker. Handcrafted runs are not drawn as curves.

tighten_y: when true, derive a bounded probability range from the displayed mean curves while retaining 0.5 as a visual reference. The default export remains the full [0, 1] scale.

If save_path is set, the base file is saved without an inline legend and save_legend_variant=True also writes a sibling "*_legend.png" file.

One panel, one curve PER ALGORITHM FAMILY, using the stronger available variant (base or CER).

Selection rule
  • Compare base and CER variants within an algorithm family.
  • Score each variant by the mean final plotted EWMA value across runs, using the same reduce_agents/agent view as the plotted curve.
  • Keep the base variant on ties so CER is only shown when it wins.

include_handcrafted: when true, add the best handcrafted-reward result as a single endpoint marker. Handcrafted runs do not participate in family selection and are not drawn as curves.

tighten_y: when true, derive a bounded probability range from the displayed mean curves while retaining 0.5 as a visual reference. The default export remains the full [0, 1] scale.

If save_path is set, the base file is saved without an inline legend and save_legend_variant=True also writes a sibling "*_legend.png" file.

Write suite-level algorithm comparison plots for average-agent, best-agent, and one-agent satisfaction views.

Remove only the previous suite-root average-agent comparison artifacts.

Legend Exports

Styling