Mechanical Turk had a good run, but not surprised it's shutting down. I'm sure the platform was getting flooded with people doing task arbitrage and using lots of AI anyway.
I believe the issue is that this can no longer be a horizontal play. MTurk was mostly for unskilled tasks...the kind AI can do well enough that it isn't worth the cost differential to verify it or keep farmed to humans. The "trust but verify" AI output is now the kind that requires domain expertise. This is what most full stack AI companies are bringing to industries.
Curious if this kind of work will come around again one day or was just a moment in time. If it does I'm sure it will be specifically about generating training data.
This is what I meant by full stack AI companies. I don't think you could get humans into the loop fast enough if they didn't have some idea of the type of task involved. I don't want people to be asked to fold a tshirt one moment and do a difficult traffic merge the next.
There is training systems and validation of skills in mturk iirc: for tshirt folding, you'd be given fake setups to be able to get used to controlling the robot, if you can't do it, you won't ever get assignments to do it. For traffic overrides, you'd be tested on having correct knowledge, and once again given supervised tasks to show you can actually be trusted (and there would be safety systems, elevating tasks that can't be performed at your level to people who can, etc)
The worker needs to do a bunch to opt into any given work group, which makes the (lack of) payments extremely unreasonable on top of everything else
It doesn't really matter what you want though, only what CEOs want and that's low costs. I can see a combined shirt folding/traffic merging platform taking off.
That's a remarkable idea. It could be heavily gamified, it could train models, and it might actually be mentally stimulating since you'd be facing different situations all the time.
Except, I'm a grown adult and I can't fold a t-shirt properly
It is weird because the last time I've heard about MTurk was about developing countries being rather reliant on it for doing AI grunt work. If am not totally wrong this must mean that the data work has moved to other services.
I'm unsure. If it does, it will be work that is too expensive or inaccurate or regulatory for current AI methods. For example, you want a doctor to sign off on some AI output on a diagnosis.
However, I'm not sure a single platform will be how it emerges
As AMT's largest requester for the past 10 years, this news was relayed to requesters at the same time as respondents. It's also worth noting that our lead contact, the Sr Program Manager at AWS leading AMT, transitioned to Amazon Bedrock and SageMaker Model Evaluations a ~2-3 years ago.. Leaving behind essential zero team managing the project after they migrated over the stored value accounts to native AWS billing.
Initially the annotation of biomedical literature (https://pubmed.ncbi.nlm.nih.gov/25592589/) but left academia for commercial use where we transitioned to arbitrage the market research and ephemeral task markets. It was essentially a research panel with a underdeveloped UI for tasks.. so abstract away AMT's tooling so it can be used by any buyer in the ResTech, Political Polling and Data Annotation space
My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals, their commission algorithm would round down to the nearest whole cent when their 20% commission resulted in a fractional cent on the unit transaction level, not the monthly invoice level. Was ultimately worth it to have implemented https://git.generalresearch.com/panels/amt-jb/tree/jb/flow/a...
> My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals, their commission algorithm would round down to the nearest whole cent when their 20% commission resulted in a fractional cent on the unit transaction level, not the monthly invoice level.
> Initially the annotation of biomedical literature but left academia for commercial use where we transitioned to arbitrage the market research and ephemeral task markets. It was essentially a research panel with a underdeveloped UI for tasks.. so abstract away AMT's tooling so it can be used by any buyer in the ResTech, Political Polling and Data Annotation space.
Reading this is similar to how I feel when I've asked Claude about something it coded for me. As with Claude, I think I get it after reading it three times; you guys were a middleman that provided a simplified interface for Mechanical Turk?
Buyer A is presented with the following rate card:
< 10min targeting gen pop: $3
10 <= 15min targeting gen pop: $5
Of course they'll bid saying their 14 min survey only takes 9 min to complete. Buyers try to cheat pricing strategies of exchanges just as much as respondents try to cheat buyers on survey platforms. Both parties can't be trusted and have adverse incentives.
“ General Research combines coordinated operations to identify and neutralize foreign actors using technological advancements, international covert operations, and networking analysis as an Internet Service Provider.”
> General Research Laboratories, LLC (“GRL”) is an aggregator of market research surveys in multiple marketplaces for business customers. GRL does not typically host consumer surveys which are conducted by other consumer-facing organizations....
So, just a middleman to give you fake/shady at best survey responses to pad your numbers, so big enterprises/concultancies can have data that say whatever they want. The rest of the website is just BS
Finally you get it. Only exception is that we run our own exchange now to do task bidding so we don’t need to deal exclusively with other companies to middleman. Core business is what’s called yield management (akin to DSP in adtech) where the best survey (is the user qualified for it, does it have the best pay, etc) is selected for traffic in <100ms. I’d only argue the shady companies are the ones paying proxies (like cint.com) and pushing paid user acquisition instead of surveys. We actively fund ontology development for better profiling targeting. but yes, big enterprises/consultancies can use and interpret the collected data however they want, not our responsibility and we have no legal rights over it anyway
It should be a bipartisan issue that a Swedish company is paying a UAE Residential Proxy company [1] to build tools that allow people from anywhere in the world to take US political polls that are used by both parties to collect election data.
Just for starters, you can't even think about soliciting online work from a panel without a robust residential proxy detection methods. We had to build our own:
Polling data isn't some guaranteed right or something. If you ask the Internet questions and then sell the collated answers, it's your job to authenticate the responses or your polls will be wildly wrong and people will stop buying them.
Which is already the case and nobody but a bunch of campaign contractors who can't justify their do nothing jobs anymore cares.
Obviously, a problem not even fully addressed by L2’s datasets. Sure, but the use of proxies to masquerade the identity of users being sold to researcher buyers that paid for a different service is fraud
Users/respondents lie, and many buyers/researchers are neo-luddites; most C level still come from the telemarketing days. A fun example: mobile targeting is terrible when off wifi because all survey platforms uniq identify users based off their IPv4, so any T-Mobile LTE users in the same city going through the same CGNAT often share profiling data that ends up conflicting, which ends up meaning they don't get sent into the best survey(s). The issue is even worse in heavy IPv6 countries like India+France.
So, I have a story to share about Mechanical Turk that you might find interesting. I’ve shared it a couple of times before on Twitter and Bluesky, but I’ll share it again.
Long story short: Mechanical Turk saved my bacon.
Back in 2005, I was working a job at a small-town newspaper in a town I’d never lived before. I didn’t know anyone beyond the staff (as I was working layout rather than as a reporter), and I had gotten interested in the idea of doing Mechanical Turk for a few extra bucks. I thought there might be a formative scene of interested people doing this, so I started working on a blog for it. I briefly collaborated on it with another guy who put it on forum software because he didn’t know how to use, like Drupal.
That blog was called Turking.com, and it seemed like it was going well for a bit. We even got a mention on the AWS website. But after about two or three months, it was clear the initial excitement around the idea (and the initial work) had died down. (The work picked up later, but in clearly different ways. I don’t think folks were really “excited” about it after that point.)
The site had started to die out in part because my iBook suffered a catastrophic GPU failure and me, being a broke small-town newspaper employee, did not have the money to replace it. Plus, the community just hadn’t emerged like we had expected.
But then I got an email out of the blue: Someone wanted to buy my domain, which I owned outright, but they didn’t want to say who. I got contacted by a broker, and the exchange took place over escrow.
(The domain most assuredly was bought by Amazon and is managed by MarkMonitor.)
I didn’t get a ton of money from the deal, but I did get enough to pay for a new laptop. I had some regrets about selling it (in part because I originally built the site with someone else), but I was in a bit of a dire financial situation which that proved to be the starting point for getting myself out of.
That domain purchase was notably more than I ever made from clicking and classifying random pictures, that said.
I explained the situation after the fact, and they were understanding. But certainly I admit that if I could do it again I probably would have clued them in sooner. The site had slowed down by the point this happened, so I’d describe it as a little more of a fire sale.
I learned a lot from that situation that I took to future sites.
It’s kind of crazy they’re shutting this down just when this Service probably has the most possibilities ever. You have an agent with Multiple people doing actual physical tasks in the real world seems like something that could be really powerful.
In spite of the name I think the overwhelming majority of the tasks were things that could be done by LLMs and they're probably getting flooded by people using bots to do exactly that. Even before the age of LLMs Mechanical Turk data was pretty bad because you'd have a bunch of people racing to answer questions as quickly as possible to get their $0.25 or whatever. So it was essentially a test of 'can you input random answers as quickly as possible while paying enough attention to notice the attention question that says to mark d.'
Any remotely open platform for this sort of stuff is going to wrecked by LLMs. But making some sort of high trust, high verification market is going to end up sending prices high enough that it becomes an unattractive proposition for use. And even in that case it's just going to be a cat and mouse game of people figuring out how to game the system enough to get trusted before handing it off to the LLM.
The concept behind it isn't going away. There are plenty of companies that hire people in bulk to do manual data labeling, transcription, RLHF, moderation and lots more. It's just that they are now catering to large AI companies, not regular people looking to get some repetitive work done (since that can now mostly be done by AI).
The only time I ever used it (as a turk?) was when Jim Gray went missing at sea and satellite imagery of vast regions of the pacific were fed through Mechanical Turk. Seems a trivial problem now, but was not then, and needed humans to take a look.
I was just thinking about the good old days combining human judgements on mturk for classifying high-value forum threads. Since replaced by using judgements from multiple LLMs.
The only use I can think of it these days is for social scientists who need to gather judgements from bona fide humans.
This will be devastating for a lot of workers in low-infrastructure countries.
I used it to have people transcribe my dad’s handwritten letters and journals. It was touching to get notes from a “Turk” saying how much she enjoyed his travels and following the cast of characters in his life!
Absent visibility into AWS P&Ls--even notwithstanding AI--Mechanical Turk has been a real outlier at AWS for a number of years now, so not really surprising.
Had some absolutely bizarre results from their attempt to integrate mturk with Bedrock's 'ground truth' thing as of a few months ago. Threw simple mnist digits recognition at it, see what the quality, timing, and cost was. Figured mnist digits was at this point trivial. Spent $10 and the accuracy was marginally better than guessing, completely unusable results. Was completely baffled, people were publishing peer reviewed research based on exclusively mturk results.
Surprised Amazon managed to fumble mechanical turk at the same time Mercor/Scale and all these other companies started hiring humans to do data labeling tasks.
I made a couple thousand dollars off it squeezing in tasks here and there between meetings. Amazon Payments were challenging to redeem as I recall. I did more or less write the bulk of someone’s doctoral thesis. Kept giving me a dollar to summarize the findings of various psychology papers. The most memorable was the one where people were put in a room and someone sprayed “liquid ass” on the wall. The subjects given no explanation experienced higher levels of anxiety than those that were told there was a sewage leak being repaired. One of the strangest dollars I ever made. My PS4 and game collection was spectacular.
I believe the issue is that this can no longer be a horizontal play. MTurk was mostly for unskilled tasks...the kind AI can do well enough that it isn't worth the cost differential to verify it or keep farmed to humans. The "trust but verify" AI output is now the kind that requires domain expertise. This is what most full stack AI companies are bringing to industries.
Curious if this kind of work will come around again one day or was just a moment in time. If it does I'm sure it will be specifically about generating training data.
"This robot is having trouble folding a tshirt help it out for 1$"
Unless they go the waymo route of highly trusted people but I think mass deployed robots are a bit safer than a car for this.
The worker needs to do a bunch to opt into any given work group, which makes the (lack of) payments extremely unreasonable on top of everything else
Except, I'm a grown adult and I can't fold a t-shirt properly
Oh, I remember UpWork.
However, I'm not sure a single platform will be how it emerges
My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals, their commission algorithm would round down to the nearest whole cent when their 20% commission resulted in a fractional cent on the unit transaction level, not the monthly invoice level. Was ultimately worth it to have implemented https://git.generalresearch.com/panels/amt-jb/tree/jb/flow/a...
can someone ELI5 because what in the office space
> Initially the annotation of biomedical literature but left academia for commercial use where we transitioned to arbitrage the market research and ephemeral task markets. It was essentially a research panel with a underdeveloped UI for tasks.. so abstract away AMT's tooling so it can be used by any buyer in the ResTech, Political Polling and Data Annotation space.
Reading this is similar to how I feel when I've asked Claude about something it coded for me. As with Claude, I think I get it after reading it three times; you guys were a middleman that provided a simplified interface for Mechanical Turk?
https://www.instagram.com/p/DZacltqHMjT/?img_index=8&igsi=bT...
> Ripe with fraud, labor exploitation, political polling manipulation
I am not sure if they want to convey what I think I am reading or not.
[0] https://generalresearch.com/mission/#:~:text=Our%20Customers
< 10min targeting gen pop: $3 10 <= 15min targeting gen pop: $5
Of course they'll bid saying their 14 min survey only takes 9 min to complete. Buyers try to cheat pricing strategies of exchanges just as much as respondents try to cheat buyers on survey platforms. Both parties can't be trusted and have adverse incentives.
What is so hard to understand????
> General Research Laboratories, LLC (“GRL”) is an aggregator of market research surveys in multiple marketplaces for business customers. GRL does not typically host consumer surveys which are conducted by other consumer-facing organizations....
So, just a middleman to give you fake/shady at best survey responses to pad your numbers, so big enterprises/concultancies can have data that say whatever they want. The rest of the website is just BS
> Our Customers
> Ripe with fraud, labor exploitation, political polling manipulation; our customers demand the best tools and security ...
It should be a bipartisan issue that a Swedish company is paying a UAE Residential Proxy company [1] to build tools that allow people from anywhere in the world to take US political polls that are used by both parties to collect election data.
Just for starters, you can't even think about soliciting online work from a panel without a robust residential proxy detection methods. We had to build our own:
`wget -N 'https://grip.net/files/grip-proxy-30d.mmdb'`
[1] https://www.youtube.com/watch?v=eOmeQcwSK3o flagged by Nokia Deepfield and CTRL for it's involvement in botnets
Which is already the case and nobody but a bunch of campaign contractors who can't justify their do nothing jobs anymore cares.
https://en.wikipedia.org/wiki/Simulacron-3
Spoiler: gur jbeyq va juvpu gur ynj fhccbfrqyl rkvfgf vf n fvzhyngvba, perngrq ol be sbe cbyyfgref!
Long story short: Mechanical Turk saved my bacon.
Back in 2005, I was working a job at a small-town newspaper in a town I’d never lived before. I didn’t know anyone beyond the staff (as I was working layout rather than as a reporter), and I had gotten interested in the idea of doing Mechanical Turk for a few extra bucks. I thought there might be a formative scene of interested people doing this, so I started working on a blog for it. I briefly collaborated on it with another guy who put it on forum software because he didn’t know how to use, like Drupal.
That blog was called Turking.com, and it seemed like it was going well for a bit. We even got a mention on the AWS website. But after about two or three months, it was clear the initial excitement around the idea (and the initial work) had died down. (The work picked up later, but in clearly different ways. I don’t think folks were really “excited” about it after that point.)
If you want to get an idea of it, there was one capture on the Wayback Machine: https://web.archive.org/web/20051124231722/http://www.turkin...
The site had started to die out in part because my iBook suffered a catastrophic GPU failure and me, being a broke small-town newspaper employee, did not have the money to replace it. Plus, the community just hadn’t emerged like we had expected.
But then I got an email out of the blue: Someone wanted to buy my domain, which I owned outright, but they didn’t want to say who. I got contacted by a broker, and the exchange took place over escrow.
(The domain most assuredly was bought by Amazon and is managed by MarkMonitor.)
I didn’t get a ton of money from the deal, but I did get enough to pay for a new laptop. I had some regrets about selling it (in part because I originally built the site with someone else), but I was in a bit of a dire financial situation which that proved to be the starting point for getting myself out of.
That domain purchase was notably more than I ever made from clicking and classifying random pictures, that said.
I learned a lot from that situation that I took to future sites.
Any remotely open platform for this sort of stuff is going to wrecked by LLMs. But making some sort of high trust, high verification market is going to end up sending prices high enough that it becomes an unattractive proposition for use. And even in that case it's just going to be a cat and mouse game of people figuring out how to game the system enough to get trusted before handing it off to the LLM.
Mercor has a market value of $20B basically doing the same thing but desperately trying to find workers.
They are terribly monotonous tasks.
I can see why it's shutting down if its still the same thing.
Lllms could probably do everything there without rotting out minds for basically pennies.
The only use I can think of it these days is for social scientists who need to gather judgements from bona fide humans.
This will be devastating for a lot of workers in low-infrastructure countries.
It appears now we can get along with just a single “artificial.”