An open-source developer testing ActivityBot paid people to follow the project's installation README while sharing their screens, using live observation to uncover assumptions that were invisible to the person who wrote the instructions. The exercise cost about €150 and led to changes after each session, according to the developer's account.

The testing was included in an NLnet grant application. Participants were offered €25 for an hour, recruited through Mastodon and invited to a video call. At the start of each session, the developer explained that the goal was to evaluate the first-run experience and that direct criticism was welcome.

The most important instruction was to think aloud. Participants described what they were attempting, which steps confused them, what frustrated them and what worked well. The developer took handwritten notes rather than using an automated transcription or AI system, then revised the README before observing the next participant. That sequence allowed later tests to check the changes made after earlier problems.

The method targeted a common weakness in technical documentation: authors already know the intended path. They may mentally correct a mistyped option, remember that a command needs elevated privileges or assume a restart is obvious. A clean virtual machine can reveal environmental dependencies, but it cannot by itself show how a new user interprets ambiguous language. Watching another person work through the instructions exposes those interpretation gaps.

ActivityBot's developer says the README became demonstrably easier to follow, while stopping short of calling it perfect. The supplied account does not include a formal error rate, completion-time comparison or independent usability study. The result is therefore a practical case report rather than proof that paid testing will produce the same outcome for every project.

Paying participants was also not presented as a strict requirement. The developer recommends finding real people who will verbalize their reactions even when a project cannot afford compensation. The central principle is to observe genuine users, not to simulate a collection of users through a language model. The account argues that people contribute humor, individual perspective and audible frustration that help an author judge which problems demand attention.

The developer connected the exercise to previous technical-writing work for GOV.UK, where another person reviewed drafts, removed elaborate phrasing and caught errors that automated spelling checks missed. That experience informed a broader conclusion: documentation benefits from direct human challenge because authors are poorly placed to detect all of their own assumptions.

For small open-source projects, the experiment offers a modest usability model: recruit several people, start with a fresh installation, have them narrate each decision, fix problems between sessions and compensate their time when possible. Its lesson is less about the precise €25 payment than about treating installation instructions as a product that should be tested with the people expected to use it.