BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Tools/AI Tools by Job, and the Jobs Nothing Handles Yet
ToolsAugust 5, 2026
Read · 5 min
ai tools · best ai tools

AI Tools by Job, and the Jobs Nothing Handles Yet

An opinionated selection organised by job, with the criterion behind each pick and the four jobs where no tool is good enough yet.

Key takeaways
  • Pick by job, never by tool. A shortlist organised around what you repeat every week survives model releases; a shortlist organised around brands does not.
  • The Remote Labor Index gives the clearest measure of the ceiling: 240 real freelance projects worth about 144,000 dollars, and the best agent finished 16.1 percent of them to a standard a paying client accepted.
  • That is up from 2.5 percent eight months earlier, which is fast progress and still means five out of six whole projects come back rejected.
  • Failures cluster in finishing rather than thinking: 45.6 percent were below professional quality and 35.7 percent arrived incomplete or malformed.
  • Four jobs a small business does weekly still have no tool worth paying for, and knowing which they are saves more money than any subscription choice.
  • The only scoring rule that survives contact with reality is how much of the output you ship without rewriting it.

Most lists of the best AI tools are organised around products, which is why they are worthless within a quarter. Products get bought, renamed, repriced and folded into something else. The jobs you do every week do not change at all. Organise the question around the job and the answer stays useful even when the logo on it changes.

This piece is the other half of an argument we have made before. We have already covered which handful of tools a small business can realistically keep, and the conclusion there was about attention: roughly three tools, chosen by how often you touch them. What follows is the part that a directory never prints, which is the list of jobs where the honest recommendation is to buy nothing.

What is the actual ceiling on this stuff?

Measured, not guessed, and the measurement is unusually good. The Remote Labor Index, published in October 2025 by a group of 47 researchers led by Mantas Mazeika, does something benchmarks normally avoid. Instead of asking whether a model can answer exam questions, it hands agents 240 real freelance projects that human professionals were actually paid for, then asks human evaluators whether the delivered work would have been accepted.

The projects are not toys. Scale's write up puts the combined value at 143,991 dollars, a median project value of 200 dollars and a median completion time of 11.5 hours of human work, spread across 23 domains from CAD to graphic design to data analysis. At launch, the best performing agent completed 2.5 percent of them acceptably.

By mid 2026 that number had moved a long way. The Decoder reported the leading model at 16.1 percent, with the next model at 8.3 percent and the one after at 6.3 percent. Four times the automation rate in under eight months is a real trend and it deserves to be taken seriously in both directions. Sixteen percent is remarkable progress. It is also five whole projects rejected for every one accepted.

We looked at this benchmark in more depth when it first landed, in the piece on what the Remote Labor Index measures that other benchmarks miss. The number that matters for a buying decision is not the headline rate though. It is the shape of the failures.

Why do the failures matter more than the success rate?

Because they tell you which part of your job to keep. The rejected work broke down into quality below professional standard at 45.6 percent, incomplete or malformed deliverables at 35.7 percent, technical and file integrity problems at 17.6 percent, and inconsistencies across the files of a single project at 14.8 percent.

Read that list again with a shop owner's eye. Almost none of it is about being unable to do the work. It is about not carrying the work to a finish: files that do not open, a set of images that do not match each other, a deliverable missing its last third. The Decoder's account names the same bottleneck, operating professional software and noticing subtle flaws.

Breakdown diagram showing why delivered project work is rejected, split across quality, incompleteness, broken files and inconsistency

That gives you a rule with more predictive power than any feature comparison. A machine is worth paying for on the parts of a job that are generative and self contained. You keep the parts that require consistency across many outputs, judgement about whether something is subtly wrong, and responsibility for the finish. Every job in the table below is split along exactly that line.

The jobs, and what a machine actually does within them

Costs are given as a billing shape rather than a figure on purpose. Sticker prices move every quarter and the shape of the bill is what determines whether a tool gets expensive as you grow. We checked live prices for a shortlist of named tools in the companion piece; this table is about how the money behaves.

Job you repeat weeklyWhat the machine does wellWhat stays yoursHow it bills
Writing product and page copyFirst drafts from a spec, and variations of a line you already likeThe claim, the price, anything a customer could hold you toPer seat, or per token if you build it yourself
Answering repeat customer questionsDrafting a reply from your own policies when the question is commonAnything involving money moving back, or an unhappy customerPer conversation or per seat, rising with volume
Reading reviews and messages in bulkGrouping hundreds of texts into themes and counting themDeciding which theme is worth changing the business forPer token, so it scales with how much text you feed it
Catalogue imagesBackgrounds, cleanup, mood shots, consistent cropsAny image that stands in for what the customer receivesPer image or per credit, effectively per output
Moving data between the apps you already pay forThe plumbing, once you have specified it exactlyDeciding what should trigger what, and watching it failPer task, which punishes growth rather than headcount
Note

Notice what none of the rows say. Not one of them is a whole job. Every honest row is a stage inside a job, with a human at the start defining it and a human at the end accepting it. Any tool sold as the complete version of one of these rows is selling you the 16 percent case and quietly leaving you the other five.

How do you tell a stage from a whole job?

By asking who accepts the output. A stage ends when the work passes to you or to another tool and somebody still has to say yes. A whole job ends when it reaches a customer. Almost every claim of full automation you will read is a stage described as a job, and the tell is that the description skips the acceptance step.

Run the test on a concrete case. Generating fifty product descriptions is a stage: somebody reads them, fixes the two that invented a material, and publishes. Running a product launch is a job: pricing, stock, copy, images, and an email all have to agree with each other, and the cost of one of them disagreeing is a customer ordering something that does not exist at a price you did not set. The benchmark evidence says machines are currently good at the first shape and unreliable at the second, and the failure numbers explain why. Inconsistency across the files of a single project accounted for 14.8 percent of rejections, and a launch is nothing but files that have to agree.

This also explains a pattern people find confusing, where a tool looks brilliant in a demo and disappointing in week three. Demos are stages. Week three is jobs.

What changed in eight months, and what did not?

The automation rate more than quadrupled, which is the fastest movement on any economically grounded benchmark we track. What did not change is the composition of the failures. Agents are not failing because the underlying problems became harder. They are failing on delivery, on file integrity, on noticing that the fourth of six images does not match the other five. The same shape shows up in note taking, where a meeting summary reads well and omits the decision rather than getting the discussion wrong.

That distinction matters for planning. If the failures were about capability, waiting a year would be a strategy. Because they are about finishing and consistency, the thing that improves your odds is structural rather than temporal: give the machine smaller units of work with a clear acceptance test, and keep the assembly yourself. A shop that adopts that posture gets value out of today's tools and gets more out of next year's without changing anything about how it works.

There is also a caveat in the reporting worth carrying. Only 218 of the 240 projects were evaluated for the leading model before United States government restrictions took effect. Even assuming it failed every remaining project, its rate would sit near 14.6 percent, so the headline holds. It is a reminder that these numbers come with asterisks, and that anyone quoting a benchmark at you without the asterisk has not read it.

Which jobs still have nothing worth buying?

Four, in our view, and this is a judgement formed from building commerce tooling rather than a measured statistic. Treat it as a claim to test against your own week.

Demand forecasting for a small catalogue. The maths needs history and repetition that a small shop does not have. A hundred products with eighteen months of patchy sales and two changes of supplier will produce a forecast, and the forecast will be a restatement of your own guess with more decimal places. We set out the evidence for this in the piece on what forecasting software can and cannot do for a small catalogue, and nothing since has changed the conclusion.

Pricing decisions. A model can compute a price from rules you supply. It cannot tell you what your market will bear, because that information exists nowhere in its training data and nowhere in your spreadsheet either. Tools that claim to optimise price for a small independent seller are optimising against a demand curve nobody measured.

Anything where being wrong is expensive and checking is slow. Tax treatment, customs classification, warranty and returns law. The output looks authoritative in exactly the cases where you are least equipped to spot the error, which is the worst possible combination. Use it to draft the question you take to somebody qualified.

End to end delivery of a project with many linked files. This is the benchmark result, restated for your business. A brand refresh, a full catalogue migration, a multi page site with consistent assets: these are precisely the projects where the failure modes above concentrate. The agent will produce something impressive and something inconsistent, and finding the inconsistency costs more than doing the work.

How should the work be split between models, then?

Along the line the vendors themselves draw. OpenAI's own guidance recommends a reasoning model when accuracy dominates and the problem is multi step and ambiguous, and a standard model when speed and cost dominate and the task is well defined. It says outright that most real workflows use both, with the careful model planning and the fast model executing.

For a shop that translates cleanly. One careful pass to decide how a category should be described, priced or structured. Then many cheap passes applying that decision. The expensive mistake is paying premium rates on the repetitive half, and it is easy to make because the repetitive half is where the volume is.

Card showing a two week trial method: run ten real pieces of work, count outputs shipped without edits, keep only above six

What test tells you whether to keep a tool?

Run it on ten real pieces of your own work over two weeks and count how many outputs you shipped without rewriting. Not ten demos, not ten examples from the vendor's gallery: ten things you would have had to do anyway.

Six or more shipped as written and the tool has earned its line on the bill. Three to five and you have bought a slightly faster first draft, which is worth having only if the job is high volume. Two or fewer and the tool is costing you time, because reading and repairing a wrong output is slower than starting from nothing. This is the same rewrite rate rule we applied to the named shortlist, and it is the only score that has predicted anything for us.

Two practical notes on running that test honestly. Do it on your worst inputs rather than your best, because the good cases were never the problem. And write the count down at the time, because memory of tool performance is generous in a way the bill is not.

Does the benchmark evidence transfer to a one person business?

Partly, and the caveats cut both ways. The Remote Labor Index measures agents working autonomously from a client brief with no interaction, which is harder than how anyone actually works. A person steering a model through five turns will get further than 16.1 percent on the same project.

The other caveat runs the other way and is less comfortable. The evaluation used human judges for a reason: model judges overestimated performance by roughly two and a half to three times on newer models. If you are letting a model check its own work, or letting a tool report its own success rate, you are looking at a number inflated by about that much. The only reliable evaluator in your business is you, looking at output you are about to put your name on.

What about the tools you already pay for?

Most small businesses acquire AI features without buying an AI tool, because the software they already use adds them. The invoicing app grows a summariser, the inbox grows a reply suggester, the store platform grows a description generator. These are worth more attention than the standalone products, for a reason that has nothing to do with quality: they cost nothing extra to try and they die quietly if you ignore them.

Apply the same ten task test to them before you shop for anything new. A built in feature that passes is worth two standalone tools that pass, because it carries no additional login, no separate bill and no new place for your work to live. A built in feature that fails costs you nothing to abandon, which is the opposite of a subscription you talked yourself into.

The shortlist, compressed

Choose one tool for the job you repeat most. Give it two weeks and ten real tasks. Keep it only if you shipped six outputs unedited. Do not add a second until the first has become something you reach for without deciding to. Buy nothing at all for forecasting, for pricing, for anything legal or fiscal, and for whole projects with many linked files.

That is a shorter list than any roundup will give you, and it has one property roundups lack: every item on it is a job you already do, so nothing here depends on which model is ahead this month. When the automation rate moves from 16 percent to 30 percent, the jobs stay the same and the boundary between the machine's half and yours moves a little. Watching that boundary is the whole skill.

If the job you repeat most is building and maintaining the shop itself, that is a case where generation genuinely covers a whole stage rather than a slice of one, because the output is code you can run and check immediately rather than a deliverable someone has to accept. What separates tools in that category for a merchant is whether the generated app also knows what an order is, which is where a general purpose app generator stops and a commerce one starts. Our store builder is built around that split, and the credit pricing is per generation rather than per seat, which matches how the work actually arrives.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building