CALLIMASTER / FROM IMAGE TO ACTION

The ink
remembers
the brush.

A single image. A recovered 6-DoF trajectory.
A real brush bringing it back to life.

Watch the real robot
REAL ROBOT EXECUTION01 / 05

IMAGE SUPERVISION.
PHYSICAL EVIDENCE.

6-DoFDirectly executable brush poses
+7.16 ppReal-robot clDice vs. CalliRewrite
5 test setsFour scripts + the Collected set

01 / The idea

Expert motion is rare.
Its traces are everywhere.

Expert brush trajectories are costly to capture. Calligraphy images are abundant. CalliMaster uses these visual outcomes to learn the motion that produced them.

A differentiable brush model connects position and orientation to the resulting ink, allowing image supervision to refine full 6-DoF trajectories.

6-DoF

Beyond the centerline

Full brush position and orientation

1K

Real demonstrations

Motion-captured calligraphy characters

300K

A bridge to scale

Synthetic characters with stroke geometry

2M

Images become teachers

Real images without action annotations

02 / The method

Physics connects
the dots.

A differentiable loop connects the desired image, the predicted motion, and the ink produced by brush–paper contact.

01

See the ink

Single character image

02

Recover motion

6-DoF brush trajectory

03

Model contact

Differentiable brush physics

04

Render the trace

Reconstructed ink image

VISUAL OUTCOME SUPERVISION
FIG. 01 The full framework. Limited action demonstrations become the starting point for learning from millions of visual outcomes.
I.

History-aware brush physics

A particle-based brush captures bristle deformation, friction, and contact history to model how motion shapes ink.

II.

Differentiable contact

A neural surrogate predicts brush–paper contact. Differentiable ink rendering sends image feedback back to the trajectory.

III.

Progressive learning

Training progresses from real strokes to synthetic characters, then to diverse calligraphy images without action labels.

THE CONTACT STUDY08 / 08 — Xuan
Brush trajectory Reference ink missed by the midline reconstructionIndividual source views · click to inspect

03 / In simulation

Simulation gallery

Before the ink.
Inside the simulator.

Recovered trajectories, executed in simulation across a range of characters and scripts.

EXAMPLE 01 / AI6-DoF trajectory → simulated execution

03.2 / Simulation benchmark

The difference
is in the details.

+13.27pp

Mean simulation clDice gain over CalliRewrite¹

Select the evaluation metric
CalliMaster CalliRewrite G-HTR BaselineHigher is better ↑

All methods evaluated using simulation rendering.

¹ Simulation clDice from manuscript Table II, equally averaged across all five sets: CalliMaster 0.94746; CalliRewrite 0.81472.

View the full results table +
Scores for the selected metric and evaluation setting
Download all results

04 / Real-world execution

FIVE DEMONSTRATIONS
01 / AICALLIMASTER · REAL ROBOT
Recorded execution · 3× speed

From pixels.
To paper.

The recovered motion leaves the simulator. A physical robot executes CalliMaster’s predicted brush positions and orientations, turning a calligraphy image into ink on paper.

01 Observe02 Recover03 Write

SELECT A DEMONSTRATION

All five recordings show physical execution. Videos retain the supplied 3× playback speed.

Compare the resulting ink

04.1 / The physical evidence

10 SELECTED EXAMPLES

The same ink.
Four ways to write it.

One target image, four methods, a physical brush. Compare stroke continuity, width, and character structure across four scripts and the Collected set.

Running script
09 / 10
TargetInput image
CalliMasterOur model
CalliRewriteReal robot
G-HTRReal robot
BaselineReal robot

Target: input image. All four outputs: real robot writing.

Ink bounding boxes · normalized scale. Click to enlarge.

04.2 / Real-robot benchmark

Measured
on paper.

25 samples per method.
5 samples in each of 5 test sets.
Five metrics across four methods.

CALLIMASTER / MEAN clDice

0.8229
+7.16 pp vs. strongest baseline

Stroke-skeleton topology. Higher is better.

clDice · All five setsHIGHER IS BETTER ↑
Relative viewBaseline

Means use equal weight across the five test sets. Metrics cover all 25 evaluated samples per method; the gallery shows 10 selected examples. Labels and tables report the original scores.

Compare every test set +
clDice across all five test sets. Best values in each column are highlighted.

05 / The data

From a thousand
demonstrations.
To millions of traces.

CalliTraj-1k captures real human brush motion. Synthetic geometry and image-only supervision extend these demonstrations to diverse calligraphic forms.

Explore CalliTraj-1k
01 / ACTION SUPERVISION

CalliTraj-1k

1K

Real motion-captured characters with 6-DoF poses and approximately 7,000 individual strokes.

02 / GEOMETRY SUPERVISION

Synthetic characters

300K

Stroke masks, centerlines, and writing order connect real action priors to whole-character motion.

03 / VISUAL OUTCOME SUPERVISION

Real calligraphy images

2M

Diverse calligraphic forms supervise action learning through the ink they leave behind.

06 / Resources

Explore CalliMaster.
From the code to the ink.

CalliMaster / Abstract

The Ink Remembers
the Brush.

Recovering 6-DoF Brush Trajectories from Calligraphy Images with Visual Outcome Supervision

Learning fine-grained embodied skills is constrained by the cost of collecting expert actions, even as their visual outcomes are widely available. CalliMaster studies this opportunity through brush calligraphy, a skill shaped by deformable contact and motion history.

Our framework combines a particle-based brush simulator, a differentiable neural contact surrogate, and ink rendering into an image–trajectory–ink loop. Starting with real action priors from CalliTraj-1k, learning progresses through 300,000 synthetic characters to 2 million calligraphy images without action annotations.

A single model recovers full 6-DoF brush trajectories from a character image. Although action demonstrations contain only regular script, visual post-training extends the model to running script, cursive script, and clerical script while retaining plausible writing order. Reconstructed ink improves across the reported evaluation splits. Real-robot experiments confirm that the predicted 6-DoF poses are executable and reproduce the target ink, with a mean clDice of 0.8229.

A closer look