Tech Brew Ride Home - Wed. 11/27 – OpenAI Suspends Sora
Episode Date: November 27, 2024OpenAI has suspended access to Sora after an activist stunt. Anyone can train AI on your Bluesky posts, but that is by design, in a way. Elon is readying a straight ChatGPT competitor. And, of course,... the Weekend Longreads Suggestions. Links: OpenAI’s Sora video generator appears to have leaked (TechCrunch) OpenAI hits pause on video model Sora after artists leak access in protest (Washington Post) Someone Made a Dataset of One Million Bluesky Posts for 'Machine Learning Research' (404Media) Inside Elon Musk’s Quest to Beat OpenAI at Its Own Game (WSJ) Yes, That Viral LinkedIn Post You Read Was Probably AI-Generated (Wired) Weekend Longreads Suggestions: Should You Still Learn to Code in an A.I. World? (NYTimes) Bad influence (The Verge) Learn more about your ad choices. Visit megaphone.fm/adchoices
Transcript
Discussion (0)
On April 4th, 2023, around 2 in the morning, a man was found stabbed multiple times on a sidewalk in downtown San Francisco.
Hey, who did this to you?
What happened next turned the story into a political firestorm.
Reports have identified the victim as Bob Lee, the founder of Cash App.
From Bloomberg Podcasts, this is Foundering, the Killing of Bob Lee, beginning April 16.
Welcome to the Tech meme right home for Wednesday, November 27th, 2024. I'm Brian McCullough today. OpenAI has suspended access to SORA after an activist stunt. Anyone can train AI on your blue sky post, but that is by design in a way. Elon is reading a straight chat GPT competitor and, of course, the weekend long-read suggestions. Here's what you missed today in the world of tech.
Sora, that AI video generating tool is down. Nobody can use it. Why? Because Open AI has taken it down.
Why? Apparently, it is in response to a group of artists leaking access to the tool in protest of the company's treatment of creative professionals. In other words, a protest has taken it down. First, here's the details on the protest from TechCrunch. On Tuesday, the group published a project on the AI dev platform hugging face, seemingly connected to OpenAI's SORA API, which isn't yet publicly available. Using their authentication tokens, presumably from an early access system, the group created a front end that lets users generate videos with SORAAPI's.
SORA. Through the group's front end, any user can generate 10-second videos of up to 1080P resolution
by typing a short text description. When TechCrunch tried, the queue was quite long, but several
users on X managed to upload samples, most of which bore OpenAI's distinctive visual watermark.
As of 12.01 p.m. Eastern, the front end was no longer working. We'd venture to guess that OpenAI
and or Hugging Face revoked access. The group claims that after three hours, OpenAI shut down
Sora's early access temporarily for all artists. So why did the group?
do this. It claims that OpenAI is pressuring Sora's early testers, including red teamers and
creative partners, to spin a positive narrative around Sora and failing to fairly compensate them
for their work. Hundreds of artists provide unpaid labor through bug testing, feedback, and
experimental work for the Sora Early Access Program for a $150 billion valued company,
the group, which calls itself SORA PR Puppets, wrote in a post attached to the front end.
This Early Access program appears to be less about creative expression and critique and more about PR and advertisement.
The group didn't originally identify its members, but over the course of the day, it began to list out a few in the attachment on Hugging Face and a separate petition.
The group also claims that OpenAI is being misleading about Sora's capabilities by keeping Early Access users on a tight leash.
Every Sora output needs to be approved by OpenAI before it's shared widely, the group says, and only a few creators in the program will be selected to have their Sora created work.
screened. We are not against the use of AI technology as a tool for the arts. If we were,
we probably wouldn't have been invited to this program, the group wrote. What we don't agree with
is how this artist program has been rolled out and how the tool is shaping up ahead of a
possible public release. We are sharing this to the world in the hopes that OpenAI becomes
more open, more artist-friendly, and supports the arts beyond PR stunts, end quote. And from the
post, the fallout, quote, Open AI spokesperson Nico Felix said the company is,
is temporarily pausing all user access to SORA while it looks into the situation.
Hundreds of artists in our Alpha have shaped SORES development,
helping prioritize new features and safeguards,
Felix wrote in a statement from OpenAI.
Participation is voluntary with no obligation to provide feedback or use the tool, end quote.
So more of a creative protest than a hack,
almost an art world-style stunt.
Rather than breaching open AI systems or exposing confidential data,
they simply shared their authorized testing capabilities more broadly than intended.
Typically, AI companies maintain tight control over external testing, carefully selecting their
early access partners and often requiring NDAs or content approval processes before any generated
outputs can be shared publicly.
It's an approach that keeps feedback channels highly curated and criticism largely behind closed doors.
While this controlled testing methodology known in cybersecurity circles as red-teaming has become
an industry standard, even adopted by government agencies, there's growing pushback from
both security experts and creative professionals.
The critique centers on how these restrictive practices may actually hinder genuine oversight,
limit independent analysis, and shield companies from meaningful public accountability.
In many ways, this incident highlights the tension between controlled development and
transparent innovation in AI, at least for policing the content produced.
On the topic of AI ethics, did you know that anyone can scrape anything posted on Blue Sky and
use it to train AI? That's because their API is open by default.
Blue Sky was forced to address this after a Hugging Face employee posted a dataset of 1 million
posts from its API for Machine Learning Research.
Quoting 404 Media.
The data isn't anonymous.
In the data set, each post is listed alongside the user's decentralized identifier or DID.
Van Streen also made a search tool for finding users based on their DID and published it on Hugging
Face.
A quick skim through the first few hundred of the million posts shows people doing normal
types of blue sky posting, arguing about politics, talking about concerts, saying stuff like
the cat is gay, and wins the last time y'all had Boston baked beans. But the data set has also
swept up a lot of adult content, too. It's also noteworthy that it's a snapshot of time on
blue sky, meaning it could and probably does include since deleted posts. This data set could
be used for training and testing language models on social media content, analyzing social media
posting patterns, studying conversation structures, and reply networks, research on social
media content moderation and natural language processing tasks using social media data. The project page
says out of scope use includes building automated posting systems for Blue Sky, creating fake or
impersonated content, extracting personal information about users, and any purpose that violates
blue sky's terms of service. The data set is already popular as of writing. It's one of the top
trending hugging face projects, end quote. Blue Sky sent this statement to 404, quote,
Blue Sky is an open and public social network, much like websites on the internet itself.
Just as robots.txT files don't always prevent outside companies from crawling those sites.
The same applies here. We'd like to find a way for Blue Sky users to communicate to outside
org's developers, whether they'd consent to this, and that outside orgs respect user consent,
and we're actively discussing how to achieve this, end quote. That's the thing. Some people were
mad at X for explicitly training on user posts, but the whole thing about being at least in part
open source and decentralized, as Blue Sky purports to be, is that means all of your content is just
out there.
Sources are telling the journal that XAI plans to launch a consumer app as soon as next month.
The Elon Musk-run startup built its Colossus Data Center, which houses 100,000 GPUs in just
122 days, which again is blowing people's minds.
quoting the journal. Investors have bought into Musk's vision for XAI, or at least his record of success.
The startup has raised at least $11 billion and increased its valuation to $50 billion in a new funding round this month,
making it the second most valuable private AI developer behind OpenAI. As a money-making venture,
though, X-AI barely registers. The startup told investors its revenue is on pace to surpass $100 million annually.
Open AI expects to bring in nearly $4 billion of revenue this year. Most of XAI's revenue has
come from Musk's own web of companies. XAI's main product, its Grock Chatbot, is available only
to the subscribers of his social network X. The startup is powering customer support features for
SpaceX's Starlink Internet service, people with knowledge of the matter said. It is also expected
to help create new AI features for X's search engine, one of the people said. The startup has
discussed a deal with Tesla, whereby XAI would get some Tesla revenue in exchange for providing
the carmaker with access to its technology and resources. Now XAI is trying to stand on a
own. Earlier this month, it released a paid tool developers can use to build products using
GROC, offering discounts as an incentive. As soon as next month, it plans to launch a standalone
consumer app like ChatGBTGT, according to people familiar with the matter. In pitches to potential
employees and investors, Musk's team has touted two advantages in the race to build the most powerful
AI. One is exclusive data from X and Tesla, being used to train XAI's models. Second is an
obsessive focus on building bigger data centers faster than his competitors. The one
One in Memphis, Tennessee, dubbed Colossus was constructed in 122 days and uses 100,000 graphic
processing units or GPUs from Nvidia, making it one of the largest clusters of chips to develop
and run AI technology in the world.
XAI has told investors that it will use some of the $5 billion it raised in this month's
funding round to double the number of chips in Colossus and that it plans to raise more
money next year, according to people familiar with the matter, end quote.
An analysis suggests that over 54% of longer English language posts on LinkedIn are likely AI generated.
LinkedIn says it doesn't track how many posts are created by AI.
Quoting Wired, the Microsoft-owned social media site for business professionals has embraced AI,
even offering LinkedIn premium subscribers access to its own in-house AI writing tools that can rewrite posts,
profiles, and direct messages.
The initiative appears to be working.
Over 54% of longer English language posts on LinkedIn are likely AI-
generated, according to a new analysis shared exclusively with Wired by the AI Detection Startup
Originality AI. It's just that the corporate-speak style of AI writing on the platform can be
tricky to distinguish from genuine human-penned thought leader blogging.
Originality scanned a sample of 8,795 public LinkedIn posts over 100 words long that were
published from January 2018 to October 2024. For the first few years, the use of AI writing
tools on LinkedIn was negligible. A major increase.
then occurred at the beginning of 2023. The uptick happened when chat GPT came out, says
originality CEO John Gillum. At that point, originality found the number of likely AI-generated
posts had spiked 189 percent. It has since leveled off. LinkedIn users who spoke to Wired say
that they rely more on general-purpose large-language models to cobble their LinkedIn
posts together rather than bothering with specialty AI tools. Content writer Aletano Sebastian
says she uses Anthropics Claude to spin-off
rough drafts of posts she creates on behalf of clients in the tech industry. Of course, there's a lot of
editing done after, she says, but the chatbot still helps me save a lot of time, end quote. Something,
something. I refer you to my question yesterday about my AI avatar experiment. Time for the week on
long read suggestions. First up, according to the New York Times, U.S. coding boot camp graduates are facing
a tough job market due to AI coding tools and recent mass layoffs. According to comp,
Tia. Developer job listings are down 56% since 2019. Quote, between the time Mr. Rendon applied for the
coding boot camp and the time he graduated, what Mr. Rendon imagined as a golden ticket to a better
life had expired. About 135,000 startup and tech industry workers were laid off from their
jobs, according to one count. At the same time, new artificial intelligence tools like
chatGPT and online chatbot from OpenAI, which could be used as coding assistance,
were quickly becoming mainstream, and the outlook for coding jobs was shift.
Mr. Rennon says he didn't land a single interview. Coding boot camp graduates across the country
are facing a similarly tough job market. In Philadelphia, Mal Durham, a lawyer who wanted to change
careers, was about halfway through a part-time coding boot camp late last year, when its organizers
with the nonprofit launch code delivered disappointing news. They said, here is what the hiring
metrics look like. Things are down. The number of opportunities is down, she said. It was really
disconcerting. In Boston, Dan Pickett, the founder of a boot camp called Launch Academy, decided in May to pause his courses indefinitely because his job placement rates once as high as 90% had dwindled to below 60%. I loved what we're doing, he said. We served the market. We changed a lot of lives. The team didn't want that to turn sour, end quote. Compared with five years ago, the number of active job postings for software developers has dropped 56%, according to data compiled by Comptia. For inexperienced developers,
the plunges in even worse 67%. I would say this is the worst environment for entry-level jobs in
tech, period that I've seen in 25 years, said Venki Ganeson, a partner at the venture capital
firm Menlo Ventures, end quote. And from the verge, a look at a lawsuit that has the potential
to completely upend the influencer industry, quote, Sheal runs what is essentially a one-woman marketing
operation, making product recommendations, trying on outfits, and convincing people to buy things
they often don't really need. Every time someone purchases something using her affiliate link,
she gets a kickback. Shopping influencers like her have figured out how to build a career off of
someone else's impulse buys. But all of this, the videos, the big house, her earnings could come
crashing down. Sheal is currently embroiled in a court case centered on the very content that is
her livelihood, a Texas lawsuit in which she is being sued for damages that could reach into the
millions. In her lawsuit, Gifford alleges that Sheel copied her down to the specific frames in
videos. She claims that repeated pattern and Sheel's uncannily similar content ultimately cut into Gifford's
own earnings. The similarities extend in Gifford's telling, beyond just video content to eerie,
real-life aspects like her manner of speaking, appearance, and even tattoos. Sheal and Gifford are but
two among the many influencers making money through Amazon's program, but their case could have
paradigm-shifting consequences for everyone else. Gifford is suing Schill for a litany of offenses
stemming from what she sees as the two women's strikingly similar videos and photos on social media.
The case has potentially wide-reaching implications for influencers and creators, but it stems from a
familiar, even ordinary complaint. Gifford says Sheal won't stop copying her. In a complaint filed
in the Western District of Texas this spring, Gifford accuses Sheel of willful, intentional,
and purposeful copyright infringement in dozens of posts across platforms like TikTok and Instagram.
Gifford says there's been a pattern of copying days or weeks after she would she would
share photos or videos promoting an Amazon product, Sheal shared her own content doing the same thing.
In dozens of cases, Gifford says the angled tone or the text on Sheel's posts ripped off
hers. Exhibits submitted in the court include nearly 70 pages of side-by-side screenshots
collected by Gifford comparing her social media posts, personal website, and other platforms
where she says she'll copied her. In one instance, Gifford promoted gold earrings in the shape of a bow,
modeling them by gently swooping her hair back to show them off. Just a few days later,
Schill posted her own photos of the same earrings, similarly photographed. In another example submitted to the
court, Gifford unboxes and then tries on a white two-piece top and short set. A few weeks later,
Sheel did the same. The pattern continued for around a year, Gifford alleges. If Gifford's legal
argument is successful, it could mean any influencer making content in an established genre
could be liable, even though in general copyright law limits liability for use of genre tropes.
The really hard part for the plaintiffs in this case is to prove that in these photos and videos,
there is something protectable by copyright, that there is creativity going on here that was copied,
says Blake Reed, Associate Professor of Law at the University of Colorado Boulder. The photos in question
are relatively banal. Images of a figure wearing generic clothing, a shot of a desk with a chair
tucked in halfway. Shields lawyers argue that the imagery Gifford claims was ripped off is actually
just standard fare for influencer content that reappears again and again and which nobody can lay
claim to. It's the Amazon Hall equivalent of swinging saloon doors in a country-western film,
Reid explains. Reed says the outcome of Gifford's lawsuit will depend on whether a judge or jury
takes influencer content seriously as a creative endeavor. On one hand, it could be framed as
low-value commercial content that all looks the same, in which case Gifford's lawsuit could be
seen as an attempt to lay claim to a template of mass-produced marketing, something that copyright law
isn't really for. But a judge might see influencer content as having enough creative
weight to merit bringing copyright law into the picture. It depends a lot on what judge lands this
and how they perceive it and how it gets framed in the litigation, he says, end quote. All right, as
mentioned, it's Thanksgiving tomorrow here in the U.S., so I won't talk to you again until Monday.
I will leave you with two bonus episodes, though, from Rad History. I tried to pick two that had a
tech or business angle, so tomorrow listen for the history of the answering machine. There's
tons of stuff about tech regulation and monopolies and stuff. It's deeper than just, did you leave a
joke on your outgoing message? Believe me. And the one on Friday will be the one Farhad Manju
and I did last week on the Challenger disaster. Again, lots of tech and science there, so enjoy those.
And if you check out the rad history feed, the episode that dropped today is on the TV show
America's Funniest Home Videos. Why? Well, think of the content that was on that show,
user-generated content. I argue it was the first reality show, but also YouTube before
YouTube, TikTok before TikTok in a world of social media. Is it all America's funniest home
videos all the way down? Anyway, happy Thanksgiving to those who celebrate. Talk to you on Monday.
