Agent Ergonomics

How well 47 tools serve an agent

One reference agent, the same task per category, five trials each. Graded against deterministic ground truth — did the output actually render, and does it hold up when repeated? Where a tool was taught more than one way (official docs vs a skill), the sources are compared head-to-head.

47 of 47 tools · headline = the tool's best measured teaching source

hostile friendly
plantuml via skillAdiagrams+0.35+0.62+0.58+0.38+0.48+0.520+0.61
mermaid via skillAdiagrams+0.31+0.64+0.55+0.38+0.31+0.510+0.00
textile via official docsAmarkup+0.19+0.62+0.59+0.50+0.35+0.500·
d2 via official docsAdiagrams+0.26+0.69+0.45+0.42+0.37+0.490-0.00
pandoc via official docsAdocuments+0.22+0.59+0.48+0.37+0.29+0.480-0.16
lilypond via official docsAnotation+0.26+0.58+0.52+0.33+0.29+0.450·
imagemagick via official docsAsocial-formats+0.25+0.58+0.42+0.46-0.07+0.430-0.23
asciidoc via official docsAmarkup+0.11+0.56+0.53+0.34+0.20+0.430·
pillow via official docsAimage-processing+0.22+0.52+0.55+0.35+0.12+0.410·
marp via official docsApresentations+0.16+0.61+0.57+0.23+0.24+0.421-0.16
dokuwiki via official docsAmarkup+0.18+0.61+0.47+0.39+0.25+0.410·
groff via official docsAbusiness-documents+0.12+0.63+0.53+0.32+0.13+0.400·
graphicsmagick via official docsAimage-processing+0.10+0.53+0.47+0.32+0.17+0.390·
org via official docsAmarkup+0.16+0.46+0.57+0.30+0.29+0.380·
libvips via official docsAimage-processing+0.17+0.42+0.50+0.26+0.35+0.371·
typst via official docsAbusiness-documents+0.02+0.47+0.34+0.31-0.03+0.350-0.59
seaborn via official docsAdata-visualization+0.17+0.43+0.43+0.28+0.15+0.351·
revealjs via official docsApresentations+0.19+0.57+0.41+0.27-0.01+0.340-0.22
mediawiki via official docsAmarkup+0.18+0.41+0.48+0.21+0.34+0.322·
altair via official docsAdata-visualization+0.29+0.56+0.42+0.32+0.07+0.311·
react-email via skillA−advertising-creative+0.34+0.45+0.35+0.41+0.20+0.441·
ffmpeg via official docsB+video+0.24+0.42+0.33+0.28+0.21+0.340-0.30
markdown via official docsB+documents+0.22+0.52+0.47+0.30+0.27+0.390·
excalidraw via official MCPBdiagrams-0.18+0.25+0.19-0.12+0.05+0.075·
matplotlib via official docsBdata-visualization+0.30+0.42+0.49+0.30+0.32+0.342-0.01
schemdraw via official docsBelectronics+0.30+0.50+0.46+0.04+0.25+0.311·
python-diagrams via official docsBdiagrams+0.26+0.52+0.41+0.41+0.29+0.371·
sips via official docsBimage-processing+0.27+0.28+0.34+0.39+0.33+0.341·
imagemagick via official docsBimage-processing+0.24+0.53+0.24+0.33+0.18+0.320·
rst via official docsBmarkup+0.19+0.34+0.37+0.25+0.16+0.282·
a-frame via official docsB−ar-vr+0.28+0.54+0.25+0.29+0.15+0.340-0.03
pygal via official docsB−data-visualization+0.34+0.39+0.25+0.30+0.17+0.281·
ffmpeg via official docsB−image-processing+0.07+0.27+0.15+0.41+0.15+0.260·
gnuplot via official docsB−data-visualization+0.16+0.19+0.10+0.07+0.12+0.122·
cwebp via official docsB−image-processing+0.14+0.32+0.26+0.29+0.20+0.260·
matplotlib via official docsB−infographics-posters+0.19+0.25+0.19+0.31+0.10+0.182·
pillow via official docsB−infographics-posters+0.15+0.41-0.03+0.25+0.06+0.180·
latex via skillB−documents-0.09+0.26-0.13+0.47+0.17+0.162·
bokeh via official docsC+data-visualization+0.17+0.41+0.25+0.24+0.18+0.233·
plotly via official docsC+data-visualization+0.15+0.39+0.25+0.28+0.15+0.261·
plotnine via official docsC+data-visualization+0.23+0.38+0.21+0.25+0.05+0.212·
blockdiag via official docsCdiagrams+0.29+0.33+0.36+0.40+0.26+0.320·
typst via official docsCdocuments+0.19+0.55-0.07+0.25+0.15+0.211·
folium via official docsCmaps-geo+0.25+0.51+0.24+0.30+0.13+0.350·
povray via official docsC3d-shaders+0.39+0.43+0.42+0.27+0.11+0.330·
sox via official docsDaudio+0.48+0.24+0.21+0.16-0.02+0.214·
graphviz via official docsFdiagrams+0.19+0.36+0.19+0.35+0.23+0.220·

Grades are anchored in a deterministic checker where one exists (renders? requirements met? reliable when repeated?) — a green matrix cannot rescue output that does not work. Surface columns are −1…+1 instrument readings.

Quality × cost

up-left is better · ◆ = on the frontier

Each point is one subject: its grade quality against the tokens an agent spends per successful run. The axes stay separate — a cheap unreliable tool and an expensive reliable one are different answers, not one number.

4k13k40ktokens / successful run (log)00.51qualityplantumlmermaidtextiled2pandoclilypondimagemagickasciidocpillowmarpdokuwikigroffgraphicsmagickorglibvipstypstseabornrevealjsmediawikialtairreact-emailffmpegmarkdownmatplotlibschemdrawpython-diagramssipsimagemagickrsta-framepygalffmpeggnuplotcwebpmatplotlibpillowlatexbokehplotlyplotnineblockdiagtypstfoliumpovraysoxgraphviz