media coverage
here's the best writeups we have so far.
matt's meltdown:
Last week, Matt Mullenweg made a series of unfortunate and baffling decisions, including some of the most egregious lapses in judgment from
tumblr's AI deals:
Internal documents obtained by 404 Media show that Tumblr staff compiled users' data as part of a deal with Midjourney and OpenAI.
that second article requires an account, but the full article is on tumblr here.
importantly, the initial data dump for OpenAI included a bunch of content that â on top of all public content between 2014 and 2023 â automattic didn't mean to share, like private posts, posts on suspended or deleted blogs, unanswered or private asks, and posts that are marked NSFW.
automattic got the IDs for these posts after providing them, and requested they be removed. we don't know for sure that they have been.
notably, the article names andrew spittle, the current head of AI. he was my boss's boss's boss's boss (you get the point) before that, and openly says that he doesn't believe in work-life balance, just in doing what the job requires him to do, if that tells you anything.
AI/LLM Q&A
poking a hornet's nest here but answering a few questions i've been posed.
part of what makes this bad is that automattic already sent the data to at least OpenAI if not midjourney. and that they accidentally sent a bunch of shit they shouldn't, like the contents of deleted blogs (trust and safety still has access to those. they are not fully purged)
i do not trust that these partners would delete data retroactively when asked. if they've already fed it in for training, there's no way to undo that.
you should be able to opt out in the future, but i wouldn't trust tumblr to be able to safely truncate reblog chains and not provide that data that way. i mean, andrew spittle clearly made a call that he did not have a duty of care to anyone who took the terms of service in good faith, so i would not trust him on anything in the future.
in terms of public posts, art or otherwise, the difference is between "your stuff is maybe being scraped if it's publicly available" and "your stuff is definitely being scraped if it's publicly available". the private stuff is a bigger deal bc that should have in no way come with the expectation that someone could have fed it to AI. private asks, blogs, and deleted content should never have been available.
tumblr in general has a pretty poor understanding of how LLMs and other generative AI work. in particular art/image AI generators do not store a copy of your image/trace it/paste it into things. that would be way, way too much data for anyone to install and use. it learns from it, just like a human artist learns by looking at other people's art how color and edges work.
your work being fed into generative AI does not in any way make it more likely (especially if you are not an extremely prolific person with an extremely hard to replicate style) that people will personally bypass you as an artist to use AI to replicate your stuff. incredibly few people are good, popular, and important enough that that is an actual risk.
AI generators are a threat to creative industries as a whole in the same way machine translation threatened translators. it is something that cannot be solved by handwringing about how special your art or your feelings are. collective action (or, you know, ideally, completely destroying the current system) is the only way to prevent capitalism from adopting every new technology it can specifically in ways that involve treating workers badly.
we have a misinformation problem. we have a problem where google is getting less and less useful every day, because their algorithms are tweaked to make them money and not to be a good search engine, and because people who are paid to get useless shit to the top of the results have way more time to work on that than people writing actual useful articles.
we have a problem where workers in general are not human to the people who call the shots, and any exceptionalism about your specific art misses the point. i get it. your art matters to you. but this is not an individual issue. technology is not the problem and copyright is not the solution. generative AI and copyright both serve the interests of capital and no fight between them is a true fight that any of us benefit from.
I'm still in awe of how this has been done in semi-secret even within Automattic. The data dump for openAi seems to have been done without even informing Tumblr backend engineers. If that's not a clear signal of Spittle knowing how scummy this deal is, I don't know what it could be.



















