Agent Ergonomics
How well 47 tools serve an agent
One reference agent, the same task per category, five trials each. Graded against deterministic ground truth — did the output actually render, and does it hold up when repeated? Where a tool was taught more than one way (official docs vs a skill), the sources are compared head-to-head.
47 of 47 tools · headline = the tool's best measured teaching source
| plantuml via skill | A | diagrams | +0.35 | +0.62 | +0.58 | +0.38 | +0.48 | +0.52 | 0 | +0.61 |
| mermaid via skill | A | diagrams | +0.31 | +0.64 | +0.55 | +0.38 | +0.31 | +0.51 | 0 | +0.00 |
| textile via official docs | A | markup | +0.19 | +0.62 | +0.59 | +0.50 | +0.35 | +0.50 | 0 | · |
| d2 via official docs | A | diagrams | +0.26 | +0.69 | +0.45 | +0.42 | +0.37 | +0.49 | 0 | -0.00 |
| pandoc via official docs | A | documents | +0.22 | +0.59 | +0.48 | +0.37 | +0.29 | +0.48 | 0 | -0.16 |
| lilypond via official docs | A | notation | +0.26 | +0.58 | +0.52 | +0.33 | +0.29 | +0.45 | 0 | · |
| imagemagick via official docs | A | social-formats | +0.25 | +0.58 | +0.42 | +0.46 | -0.07 | +0.43 | 0 | -0.23 |
| asciidoc via official docs | A | markup | +0.11 | +0.56 | +0.53 | +0.34 | +0.20 | +0.43 | 0 | · |
| pillow via official docs | A | image-processing | +0.22 | +0.52 | +0.55 | +0.35 | +0.12 | +0.41 | 0 | · |
| marp via official docs | A | presentations | +0.16 | +0.61 | +0.57 | +0.23 | +0.24 | +0.42 | 1 | -0.16 |
| dokuwiki via official docs | A | markup | +0.18 | +0.61 | +0.47 | +0.39 | +0.25 | +0.41 | 0 | · |
| groff via official docs | A | business-documents | +0.12 | +0.63 | +0.53 | +0.32 | +0.13 | +0.40 | 0 | · |
| graphicsmagick via official docs | A | image-processing | +0.10 | +0.53 | +0.47 | +0.32 | +0.17 | +0.39 | 0 | · |
| org via official docs | A | markup | +0.16 | +0.46 | +0.57 | +0.30 | +0.29 | +0.38 | 0 | · |
| libvips via official docs | A | image-processing | +0.17 | +0.42 | +0.50 | +0.26 | +0.35 | +0.37 | 1 | · |
| typst via official docs | A | business-documents | +0.02 | +0.47 | +0.34 | +0.31 | -0.03 | +0.35 | 0 | -0.59 |
| seaborn via official docs | A | data-visualization | +0.17 | +0.43 | +0.43 | +0.28 | +0.15 | +0.35 | 1 | · |
| revealjs via official docs | A | presentations | +0.19 | +0.57 | +0.41 | +0.27 | -0.01 | +0.34 | 0 | -0.22 |
| mediawiki via official docs | A | markup | +0.18 | +0.41 | +0.48 | +0.21 | +0.34 | +0.32 | 2 | · |
| altair via official docs | A | data-visualization | +0.29 | +0.56 | +0.42 | +0.32 | +0.07 | +0.31 | 1 | · |
| react-email via skill | A− | advertising-creative | +0.34 | +0.45 | +0.35 | +0.41 | +0.20 | +0.44 | 1 | · |
| ffmpeg via official docs | B+ | video | +0.24 | +0.42 | +0.33 | +0.28 | +0.21 | +0.34 | 0 | -0.30 |
| markdown via official docs | B+ | documents | +0.22 | +0.52 | +0.47 | +0.30 | +0.27 | +0.39 | 0 | · |
| excalidraw via official MCP | B | diagrams | -0.18 | +0.25 | +0.19 | -0.12 | +0.05 | +0.07 | 5 | · |
| matplotlib via official docs | B | data-visualization | +0.30 | +0.42 | +0.49 | +0.30 | +0.32 | +0.34 | 2 | -0.01 |
| schemdraw via official docs | B | electronics | +0.30 | +0.50 | +0.46 | +0.04 | +0.25 | +0.31 | 1 | · |
| python-diagrams via official docs | B | diagrams | +0.26 | +0.52 | +0.41 | +0.41 | +0.29 | +0.37 | 1 | · |
| sips via official docs | B | image-processing | +0.27 | +0.28 | +0.34 | +0.39 | +0.33 | +0.34 | 1 | · |
| imagemagick via official docs | B | image-processing | +0.24 | +0.53 | +0.24 | +0.33 | +0.18 | +0.32 | 0 | · |
| rst via official docs | B | markup | +0.19 | +0.34 | +0.37 | +0.25 | +0.16 | +0.28 | 2 | · |
| a-frame via official docs | B− | ar-vr | +0.28 | +0.54 | +0.25 | +0.29 | +0.15 | +0.34 | 0 | -0.03 |
| pygal via official docs | B− | data-visualization | +0.34 | +0.39 | +0.25 | +0.30 | +0.17 | +0.28 | 1 | · |
| ffmpeg via official docs | B− | image-processing | +0.07 | +0.27 | +0.15 | +0.41 | +0.15 | +0.26 | 0 | · |
| gnuplot via official docs | B− | data-visualization | +0.16 | +0.19 | +0.10 | +0.07 | +0.12 | +0.12 | 2 | · |
| cwebp via official docs | B− | image-processing | +0.14 | +0.32 | +0.26 | +0.29 | +0.20 | +0.26 | 0 | · |
| matplotlib via official docs | B− | infographics-posters | +0.19 | +0.25 | +0.19 | +0.31 | +0.10 | +0.18 | 2 | · |
| pillow via official docs | B− | infographics-posters | +0.15 | +0.41 | -0.03 | +0.25 | +0.06 | +0.18 | 0 | · |
| latex via skill | B− | documents | -0.09 | +0.26 | -0.13 | +0.47 | +0.17 | +0.16 | 2 | · |
| bokeh via official docs | C+ | data-visualization | +0.17 | +0.41 | +0.25 | +0.24 | +0.18 | +0.23 | 3 | · |
| plotly via official docs | C+ | data-visualization | +0.15 | +0.39 | +0.25 | +0.28 | +0.15 | +0.26 | 1 | · |
| plotnine via official docs | C+ | data-visualization | +0.23 | +0.38 | +0.21 | +0.25 | +0.05 | +0.21 | 2 | · |
| blockdiag via official docs | C | diagrams | +0.29 | +0.33 | +0.36 | +0.40 | +0.26 | +0.32 | 0 | · |
| typst via official docs | C | documents | +0.19 | +0.55 | -0.07 | +0.25 | +0.15 | +0.21 | 1 | · |
| folium via official docs | C | maps-geo | +0.25 | +0.51 | +0.24 | +0.30 | +0.13 | +0.35 | 0 | · |
| povray via official docs | C | 3d-shaders | +0.39 | +0.43 | +0.42 | +0.27 | +0.11 | +0.33 | 0 | · |
| sox via official docs | D | audio | +0.48 | +0.24 | +0.21 | +0.16 | -0.02 | +0.21 | 4 | · |
| graphviz via official docs | F | diagrams | +0.19 | +0.36 | +0.19 | +0.35 | +0.23 | +0.22 | 0 | · |
Grades are anchored in a deterministic checker where one exists (renders? requirements met? reliable when repeated?) — a green matrix cannot rescue output that does not work. Surface columns are −1…+1 instrument readings.
Quality × cost
up-left is better · ◆ = on the frontierEach point is one subject: its grade quality against the tokens an agent spends per successful run. The axes stay separate — a cheap unreliable tool and an expensive reliable one are different answers, not one number.