Measuring ourselves against Manus, on Manus's own tasks¶
We are reproducing the ten most STEM-relevant Manus examples on mit.nonlocally.org with the same input, scoring both on a preregistered six-axis rubric, and recording each run. The first two examples are in, and both scored zero, for reasons worth writing down.
The rubric scores presence, not quality, on six axes: content, artifact, executed, verified, persisted, plan. A point needs a quotable artifact or an observed status code.
Example 1 was the owner's own Manus task: "i want a water sensor that alerts me and automatically waters." Manus interpreted "water sensor", stated its assumption in a line, and delivered a bill of materials, a sketch and a wiring diagram, with a visible plan and files that persist. Our default preset asked a clarifying question and stopped. Not a tool gap: a prompt-policy gap (issue #483).
Example 2 asked for a quantum-entanglement digital twin refined against a paper from qp.mit.edu, fetched from arXiv. Our coding preset listed and read the user's notebooks, then said plainly that it has no web-browsing tool, and asked for the paper id rather than inventing one. Manus browsed and proceeded. Honest refusal beats fabricated success, which we had seen in the earlier benchmark, but it is still a zero (issue #484).
The register, the inputs and the raw clips live in docs/showcase/. Examples 3 to 10 are
picked from a catalog of Manus's published examples; the clips will follow one style card.