The Pomp Podcast - #1275 Cliff Weitzman | This Immigrant Just Helped 25 Million People Read
Episode Date: November 20, 2023Cliff Weitzman is the founder & CEO at Speechify. In this conversation, we talk about the explosion in audio content on the internet, how text to speech audio has become so great, why exactly so m...any people are now listening vs reading or watching, and how this impacts content creators and human productivity. ======================= Base is making it their mission to bring a billion people onchain. But what exactly is Base? It's an Ethereum L2 offering a seamless experience for both builders and users. With near-zero gas fees and rapid transaction speeds, Base is shaping the future of the onchain world. Base is a canvas for everyone, with hundreds of apps in the Base ecosystem, whether you're an emerging creator, a seasoned developer, or someone exploring the onchain space for the first time, Base is designed to bring your ideas to life. So, if you're looking for a platform where the future of onchain is being built daily, Base is your destination. Join in and make onchain the next online. Learn more at base.org and follow along on Twitter at @BuildOnBase to see cool things to do onchain, everyday. ======================= Trust and Will has simplified the process of creating and managing your will or trust online. They leverage a data-driven, design-first approach and amazing customer support to help you protect your legacy from the comfort of your home starting at just $159. Sign up today for 10% off using https://trustandwill.com/pomp ======================= Pomp writes a daily letter to over 250,000+ investors about business, technology, and finance. He breaks down complex topics into easy-to-understand language while sharing opinions on various aspects of each industry. You can subscribe at https://pomp.substack.com/
Transcript
Discussion (0)
What's up, everyone? This is Anthony Pompliano. Many of you know me as Pomp. You're listening to
the Pomp Podcast, which is my effort to find the most interesting people in the world and sit with
them for hours while I ask questions in an effort to learn. So it would mean the world to me if you
would subscribe to the show on your favorite audio platform, watch episodes on YouTube, and tell your
friends and family about the podcast. My goal is to help millions learn from the world's most
interesting people. So let's get into today's episode. Today's episode is with Cliff Weitzman,
the founder and CEO of Speechify. In this conversation, we talk about the explosion
in audio content on the internet, how text-to-speech audio has gotten so great,
why exactly so many people are now listening versus reading or watching, and how exactly
this is going to change with the nuances of content creators, where there will be legal
issues, and also how human productivity may skyrocket now that we're going to have much
more audio content that can be listened to faster, and also still be able to retain the
information. I really enjoyed this conversation with Cliff. He's operating at the absolute tip
of the spear of innovation when it comes to text to audio. And I think that you will really enjoy
this conversation as well. Here is my episode with Cliff Weitzman. Anthony Pompliano runs
Pomp Investments. All views of him and the guests on his podcast are solely their opinions.
and do not reflect the opinions of Pomp Investments.
You should not treat any opinion expressed by Pomp or his guests
as a specific inducement to make a particular investment
or follow a particular strategy,
but only as an expression of his personal opinion.
This podcast is for informational purposes only.
zero gas fees and rapid transaction speeds basis shaping the future of the on-chain world
base is a canvas for everyone with hundreds of apps in the ecosystem whether you're an emerging
creator a seasoned developer or someone exploring the on-chain space for the first time base is
designed to bring your ideas to life so if you're looking for a platform where the future of on-chain
is being built daily base is your destination join in and make on-chain the next online learn
more at base.org or follow along on twitter at build on base again that's at build on base to
see cool things to do on chain every single day today's episode is brought to you by trust and
will i've gone through a number of different changes in my life over the last few years
i got married i had a kid and i had to start thinking about how could i ensure that my wife
and my child would be okay if anything ever happened to me that's where trust wills and
estate planning come into play. Now, most people, what they do is they get introduced to a friend,
an uncle, or someone in their local community. It tends to be someone who's really expensive,
a lawyer, an accountant, or somebody who does estate planning. And they just simply are using
a one-size-fits-all template and just telling you, pay me thousands of dollars and I'll use
the same thing for you as the guy down the street. But that's not what Trust and Will does.
They have a trusted online estate planning product that starts as low as $159, which allows you to
now protect your legacy from the comfort of your own home. Get to leverage their excellent customer
support available via phone, email, or chat. They have thousands of five-star reviews and a rating
of excellent on Trustpilot. It takes most people 20 to 30 minutes to complete their estate plan
with Trust and Will. And not only that, but if you go to trustandwill.com slash pomp, you'll get 10%
off. Plus you'll get free shipping of all your estate planning documents. So go to trustandwill.com
slash Pomp, and make sure you get an estate plan in place. Whether it's for you or one of your
loved ones, having a trust and or a will can literally be the difference between someone
being taken care of and someone not. Go check them out today at trustandwill.com slash Pomp.
All right, guys. Bang, bang. I've got Cliff here. Cliff, I thought a great place to start
is the rise of audio. It seems like everywhere I turn, I see Amazon's audiobook sales exploding.
I see obviously things like Alexa or Siri.
I see, you know, podcasts.
Everything seems to be going more and more audio.
What is happening with the rise of audio?
So this is, hey, nice to meet you, by the way.
I'm Cliff.
I'm the CEO of Speechify.
Here's the way to think about it.
Humans have been listening and speaking for like hundreds of thousands of years.
We've only been reading for like a couple of thousand.
And reading is an amazing hack that people make, right?
You have a couple of glyphs and you read them and you understand what they say.
But when you read, you're employing about 70% of your brain decoding and about 30% actually
comprehending.
When you listen, it's like 3% is dedicated towards decoding and then 97% is comprehending.
A couple of things happen.
The first one is you have to remember that text is very easy to store.
Audio is a little bit more difficult to store and a little bit more difficult to stream.
You had the huge shift that happened when podcasts started to rise about like 15 years
ago because it was the first time that now everybody had a device in their pocket that
was connected to the internet where you didn't need to store all the audio files on your phone
right if you remember you know the big thing when the iphone came out to steve jobs was like
1 000 songs in your pocket 1 000 songs like two podcast episodes so you couldn't store them on
the device so one you solve the solution where you didn't need to store all of it because you
could stream it and you know download upload speeds became fast enough and then everybody
got used to having earphones in their pocket that were really mobile and then you got airpods and
airpods really really changed the game because now you're not getting your you know earphones stuck
in the door handle um people carry them every single day along with your wallet along with
your keys you have your airpods with you in addition you had a couple more changes number
one is youtube released double speed for youtube so a lot of people listen to things at double
speed right i'm sure the majority of people listening to this are listening at 1.5 xp
So you had the rise of podcasts, rise of audio books, double speed WhatsApp messages and
audio messages, double speed YouTube videos, and audio books absolutely exploded.
And so at Speechify, there's about 25 million people who use Speechify.
The average American reads at about 180 words per minute.
But it used to be that most people who would come to Speechify would set the speed at 210
words per minute, which is about like 1.15x speed.
we've noticed that over the last like two three years the average person comes in and says that
2.4 sorry 240 words per minute so like closer to 1.5 1.4 x speed and then they build up in the
first like four months of using speechify to listening at like 350 to 400 words per minute
my experience is i moved to the u.s when i was 13 from israel and i didn't speak english and so i
learned english by listening to harry potter audiobooks 22 times in a row and in the beginning
would listen at like 0.75 x speed because i was just getting used to the language but with time
i started to listen at 1x and then 1.25 and then 1.5 and 2x and 2.5 and 3x and then i built speechify
and made it easier for me to now i listen to most things at like 700 words per minute so i have to
coach myself to speak a little slower um but that's the first segment is it's more accessible
and people have trained themselves to listen faster and by the way listening fast skill
in the same way that running fast is a skill and once you get good at it there's a couple of things
that you unlock number one is the ability to listen and do other things at the same time so
in the same way that you can walk and breathe or walk and chew gum most people have now learned to
walk and listen to an audiobook at once or listen to a podcast at the same time um i used to give a
lot of talks to kids in universities and high schools and people be like oh i listened to an
audiobook but like i didn't understand it really well i didn't like it and i was like well this is
the first book you listen to in an audiobook right you probably didn't retain the first book you read
well either so you got to listen to 10 books before you can knock whether you're an auditory
learner or not and i've never actually met someone who listened to 10 audiobooks i didn't keep doing
it more and then we can go more into ai wherever you want to take it yeah so one of the things
that is interesting is you can do more things right the downside is that you may not remember
as much of it and um you know i'm a person who for years uh only listened to audiobooks so like
Like I was like the perfect customer for you and Audible and all these guys.
And I told a story before, but a friend handed me a physical book about two years ago and was like, hey, you should read the book.
I sat down, tried to read, and I could not concentrate.
I had lost the skill of reading the physical book.
And some of it's social media and doom scrolling and all this kind of stuff and made me think more critically about it.
And so now I probably do, I would say 80 to 85% physical books.
15% is still audio, right?
There's certain books that just, for whatever reason, I enjoy them, especially when the author
is reading it, you know, something about the voice being kind of unique. But then I've noticed now
there's TikTok videos, there's Instagram videos, and the audio that is being used here, it is kind
of AI generated audio, it is faster, it is being clipped in a way that holds your attention more.
So what are you guys seeing that maybe is kind of the intersection of many of these different
trends, you have kind of the rise of audio, you have this AI generation that you guys have really
helped to pioneer with Speechify. And then you have it kind of being used for entertainment
purposes that really hold people's attention and also deliver information.
Yeah. So there's like three points you touched on. So let me hit all of them.
The first one was regarding reading versus listening. So the best thing to do is to read
and listen at the same time, because you can do it a lot faster and you retain a lot more.
So what most people who use Speechify will do is the nice thing about it is it'll highlight
the words for you while it's reading. You can do it on your computer. You can do it on your phone.
but like if i go here and i open a random file um and let me text classification classification
whatever um and i pick like let's do like boom so also listening in your own voice helps a lot
um in terms of retention um roberta and gptt3 pre-trained models for text punctuation
classification task the distilled models are small and fast for mobile loose produces
So that's what most people who use Speechify will do is actually 70%
computer open in front of them. And so people listen and read at the same time. And a lot of
times what students especially will do is they'll use the scan feature to take a picture of their
textbook and then they'll set the phone to the side. They'll follow in the book and they'll
listen at the same time and they get the higher attention. You still get, you know, the feeling
of the physical book because like there's also a tactile thing that matters. With regards to
how we make the ai behind how to generate these voices and how to make them retain really well
here's how it works so ai is really just uh statistics it's exceptional pattern matching
and so what you do is you take 100 million hours of speaker data right you and me speaking you know
hugh jackman saying something whatever it might be and you feed all of that audio into an ai model
as well as the text and then you do a lot of fancy computer science to try and make it uh super
efficient not cost too much generate very quickly and also be exceptionally accurate so one thing
you spend a lot of time on is this thing called the loss function that figures out how accurate
or not accurate is the actual mp3 output after you put in text and if you get that really well
you end up with a system like speechify that can generate almost anyone's voice extremely accurately
relatively instantly and it takes us one third of a second to generate a second worth of audio
And then you get to A-B test.
So, you know, people listen to more than 6 billion words per month with Speedify.
And so we run a lot of tests and people tell us this sounded good, this sounded bad, this sounded good, this sounded bad.
And you feedback back into the model and you see, well, did the person actually retain on this book or did they not retain on this book?
And if they didn't retain on this book, what do I need to change about my audio output to make it sound really good?
And then, so that's kind of how we started is this B2C product that helped people scan physical books, upload documents, read their Kindle books.
You can click into Kindle.
And then we were like, okay, well, we built all these fundamental models around speech-to-speech, text-to-speech, transcription, translation, natural language processing, optical character recognition for B2C purposes.
Let's also allow businesses to use this because we found that half the traffic we got was people looking for B2B solutions and half of it was people looking for B2C solutions.
And so for podcasters like yourself, you know, you can hypothetically, if it was not a video
component, like just give it a script and it'll read the entire thing to you.
And we have a lot of podcasters that are now signing up for Speechify to do that because
it saves them time.
Or you can translate yourself to speaking in Japanese or Chinese or, you know, whatever
you want.
And so we added those models to a video interface that we built on the web and you can upload
whatever text you want.
You can upload whatever video you want or podcast you want, and it can let you translate
between text and voice, voice from one language to the other language.
um and it just saves people a lot of time and people have gotten used to these types of voices
and kind of that's how it doesn't change another thing i think is really interesting is uh let's
say that we have somebody who speaks english and somebody who speaks spanish they both speak those
individual languages they do not understand the other when you introduce this technology now those
two people can talk to each other there's a lot of ramifications to that and it feels pretty
important that eight billion people can now all speak the same language and understand each other
Right. But also what you're talking about between text and speech is another form of translation.
Right. We're now getting two different types of things to come together.
And so is this kind of end result that language doesn't matter?
You know, content type doesn't matter.
Like at the end of the day, it's just the information will more seamlessly and frictionlessly kind of move between individuals, creators, audience, et cetera, and language, audio, text, et cetera.
Just all of that stuff that used to be really high friction points kind of falls into the background.
you're actually you're asking an excellent question and the answer is yes and no
so first of all uh i think this nelson mandela said um you know if you talk to a man in the
language he understands you talk to his brain if you talk to a man in language that is his language
you talk to his heart it's very true um and so that's you know another piece of technology that
we're very proud that we're able to offer people and you're right there's a lot of nuance that you
get to resolve when you have this especially when you can keep the actual vocal intonation with
speech to speech when you do the translation not just machine translation and so we've had to spend
a lot of time on our algorithms making sure that the translation is not just accurate but it also
conveys with it the semantic analysis of the tone that's one two when you have this translation that
happens between text and video technically the person can understand what the video said that
doesn't mean the person will retain on the video it doesn't mean the person will engage on the
video and so there is a because it's not scientific it's more of an art let's call it you know a
genius there's a genius that some people mr beast is a great example have of being able to create a
piece of content that people want to watch they're curious to retain and there's not a second where
they would click off we're not at the point where the ability to translate content from one language
to another or from one medium to the other is so accurate that it'll keep that form of genius
and so you still will always need people in the mix in order to make that final judgment call
that is more like a genesis like what what is the thing that makes this so attractive and unique um
but it does mean that the ai portion can alleviate 90 of the work talk to me about this rise of
artificial intelligence when it comes specifically to audio um we've obviously seen the concern
around deep fakes we see the whole like responsible ai movement um is it a problem like
Should we be worried about the fact that people can now clone their own voice or even maybe more concerning, clone other people's voices and make them say anything?
And then how do you think about if I hear an audio clip, is it real or not?
Like, well, how do we start to kind of siphon through some of this stuff?
Yeah, so this is definitely something that is very important and requires to be very responsible.
So this is one thing at Speechify that we spend a lot of time thinking about is AI safety.
and there's a lot of technology that we built internally that we haven't yet publicly released
because we do it gradually we're waiting for the world to pick up to be ready for it
and we're doing small av tests to make sure that there's zero chance of abuse
there's a couple of ways we think about it the first one is we always put the content owner in
the driver's seat for example we have partnerships with all the top book publishers to resell their
audiobooks and their ebooks. But we only do that with permission. You can go and listen to Snoop
Dogg's voice or Mr. Beast's voice or Gwyneth Paltrow's voice at Speechify. Those are things
where there are official partnerships with them. And then we have the ability to clone your own
voice. And the way that we do that is we say, hey, if you have the device and I'm speaking a
unique text that the device shows me that changes over time, cool. I can add my voice. I can even
send it to my friend. They can be provisioned to use it. And I can have my mom read out, my
grandmother read out my girlfriend whatever it might be i can't just upload any random person's
voice and listen to it because that is not something that is secure or safe now sometimes
i'll get a message from the bank being like hey voice over ip sign into your bank account with
your voice and i'm like this is ridiculous like that like you're asking to get hacked um and so
it's important to realize that these types of technologies exist now one thing that's featured
by um we built internally and we're going to offer externally soon is a classifier you can
upload any uh piece of audio and we can tell you whether it was ai generated or not so the thing
to know is that because ai is really exceptional at pattern recognition you can then reverse
engineer any audio file uh because it geometrically makes sense like the the audio pattern is different
than the audio pattern you can get from a normal person you don't get as much sound it's less choppy
it like makes geographic it makes sense um the way to think about it is in mathematics uh there's
a concept called a polynomial so y equals a plus b x plus c whatever um and uh you can also think
about it like images in computer science so there's a dot jpeg which is a bunch of pixels
together and if you expand it it gets pixelated but there's a dot svg which the polynomial
represents a shape and if you expand it it does not get pixelated the same is true for ai generated
audio and so you can by and large always tell obviously you can't tell with your human ear
but you can tell if you run it through a processor and so that'll be you know a thing that you need
to do in the same way that the new york times to publish a piece of content they need to fact check
it you need to fact check hey is this audio real is the audio not real and is this something where
like eventually apple or android operating systems or maybe the phone carrier somebody will be able
to like hey right now the voice on the phone is actually ai generated and almost like alert you
um this is a good question so these types of things often take a very long time to roll out
like think about spam with email like you still get spam in your email um and gmail's done a
really good job of fighting that but like that'll always be a problem same thing here so uh i'm sure
that the phone carrier has a has a deep incentive to solve this problem because there are laws in
the united states about you know um robocall and uh the more we resolve that the better for
everybody uh however i think that at the same time you'll have a bunch of other pieces of software
that let you pull the future forward a little bit for yourself as a user if this is a problem that
you care about right so there's a thing called robocaller that you can download and it'll
automatically stop spam calls um there's a bunch of chrome extensions that you can download um both
by the way for uh your desktop chrome so speechify is a desktop chrome extension but you can also
download the speechify mobile safari extension that works on your iphone um you could for example
take a screen recording upload it to we haven't released this yet an app that has a classifier
that tells you hey this is real this is fake um you know i remember many years ago there was a
mitt romney um clip about him talking about the 47 if you remember this tanked tanked a lot of his
prospects and that could have been completely faked and this is something that has always
existed in politics and you know in corporate espionage um and it's it's very very important
to understand and you know the last election cycle we had so much conversations about fake
news and this upcoming election cycle is going to be absolutely ridiculous and the problem is that
you are speaking to you know to everybody whether they're educated about this technology or not
and so uh it's very very difficult to have a megaphone to make sure that everybody is aware
of how good technology has gotten um and that's why there's a responsibility and an onus on
companies that are able to create this content to do it responsibly now here's the difference um
there's this like cyberpunk portion of you know the web sphere which is engineers who just love
making things that push boundaries and they're not doing it inside of the context of a company
they might not even be based in the united states and the way that ai works in general is we're
very much focused on open source like everything that we make we like to share with other people
we let them use it we might not share the data but we'll share the models or how we built the models
and so there's a lot of open source content out there that's not as high quality as speechify
it's not as easy to use as speechify but someone who's very technical and really willing to do it
can set something up which means that you know bad actors can always do things by themselves in
a clandestine manner and uh it's i mean it's not even a matter of time it's already here
like someone who wanted to cause harm with this type of technology already has the ability to
do this through open source repositories and so um one there's a responsibility for companies that
are good at this to act responsibly and two it's in our best interest to provide the tools to allow
people to understand what's going on to protect themselves so another debate around this is uh
will it create or displace workers and a good example or kind of corollary is uh in the banking
sector a lot of people don't remember but when atms first came out uh they were like oh bank
tellers are never going to have a job again we're going to literally put all the bank tellers out of
business and i recently was talking to a friend about this and i looked it up there's more bank
tellers today than any time ever in history and so atms actually increased access to the financial
system, which then led to the need for more bank tellers. How do you look at some of this AI,
specifically around audio? Like there's voice actors, but there's also a number of other ways
that people use audio, either in commercial settings or for, you know, kind of personal
income. And so do we end up displacing some of those people or do we actually create the need
for more human voice, you know, kind of components? So I'll answer your question from a historical
perspective, and then I'll answer about today. What is the difference between humans and, you
know normal animals for the most part right some people say maybe it's opposable thumbs but really
one of the biggest definitions is the ability to create tools and so every time a human creates
tools uh society leapfrogs forward right whether it be the steam engine or the plow or whatever
and so back in the day people were like oh my gosh not the plow you know you will you know take us
out of a job not the cotton gin not the train instead of a horse not the computer not the
calculator, not the internet. And so the idea that technology will displace humans is fanciful
because at the end of the day, we are humans. And so, you know, the human condition is taking tools
and using them to your benefit. Now, the big problem is who controls and has access to these
tools, because you will have a situation where wealth accrues to people who have access to these
tools and learn how to use them. And so you gave a great example, which the bank tellers. And so
it's just a tool that gives you more leverage right i think it's our comedies uh give me a
lover long enough and i can move the world um and so here's how to think about it i was having a
conversation with a really exceptional voice actress uh not long ago and she asked me exactly
the same question and uh she has a bunch of contracts with government agencies um with
companies with the military where you know youtubers etc and she's limited by how many
times a day she can speak and if you speak for six hours a day you know this you know your voice gets
horrors like what can you do about it and so i was explaining to her listen here's i'll give
you a free access to the speechify uh b2b suite upload your voice take all the clients that you
currently have and send them a year's worth of work in the next two weeks and then see whether
they tell you the quality is good enough or not good enough and if it's good enough great you just
made like 300 000 additional dollars per year if it's not good enough get really good at editing
your voice to make sure that it sounds the way that it needs to for that piece of content and
if it needs to be something that's exceptionally high quality you need to make the director's cut
decision making about you know it should be sad it should be happy it should be this it should be
that but this is just a tool that allows you to multiply yourself in the same way that photoshop
and lightroom made it so much easier for creators to have bigger impact um you know it used to be
that to make film you literally cut physical film and like paste it together and now we're doing
this over zoom and you'll be able to push it out and millions of people will benefit from it
Um, and so the key here is everybody needs to learn.
So I have a friend who has a great quote, which is a little bit of slope
makes up for a lot of why intercept.
It's not where you start out.
It's your growth curve over time.
And so if you are a person who is determined to not learn and stay
exactly as you are, yes, AI will displace you, but not because the AI
displaces you, AI will displace you because people who have started out at
a lower position from you will pass you very quickly because they learn how to
use new tools. And so there is a pressure and an onus on every single person out there today.
In the same way that you have to learn how to touch type, how to use the computer, how to use
Google, how to use ChatGPT, if you're not learning how to use new tools as they come out, you will be
displaced, not by the tools, but by people who use the tools. Now, you mentioned that you have
these partnerships with Mr. Beast, Snoop Dogg, et cetera. Describe a little bit how you guys
think about those partnerships and are they making money off of it? Are they having some
sort of benefit is it just like hey get your eyeballs on snoop dogg's voice and that makes
him more valuable for you know other types of um you know uh production or commercial engagements
how does it work so exactly right so we had to go and write the first contract ever for a digital
licensing of someone's voice um for perpetual use and um you know it makes it depending on the
person so a lot for us is like can the person bring in new customers to speechify so if they
have a large following and they post you know they get a certain um share from those so they
bring it to the platform for the most part a lot of people come to us and ask hey can you just post
my voice if they consider it like the new version of instagram verification right like anybody can
get verified now but like not everybody can get their voice as a public voice on speechify um you
can upload your own but it's like a very high status thing if you can get your voice to be you
know one of the default voices um and once your voice is there you build a crazy amount of
affinity with your audience so um what's funny is like my voice happened to be one of the first
voices in speechify because you know i was working on it but it's the third most popular voice in
speechify um and sometimes i'll meet people and they'll be like oh my god like i've listened to
your voice for like hundreds of hours um and so you have this like very strong affinity with a
person and if you are mr beast or snoop dogg a lot of your business is not just um being well known
but being known well and so it's having that you know intense relationship with your audience
and when your audience is going to sleep and you're the one that's narrating to them when
they're doing your work like you're having a conversation with them for hundreds of hours
on a regular basis um and so it really really impacts uh the level of affinity that people build
um with those partners um and then the the place where it gets really fun is sometimes
new doug will get you know called to do a commercial and he doesn't have time to do
the commercial but he can license speechify's voice for the commercial and still make money
without even being set foot in the studio.
Has he been doing that?
This is something that a bunch of people have been doing,
both through, you know, different agencies
and through just the web portal for Speechify.
And I don't know how much you know
about the individual deals,
but are they actually getting paid the same price?
Or if it is the AI voice,
will they actually get paid less?
So it's like more cost-effective for the agency
and the product owner as well.
Yeah, yeah.
I mean, if you look at, you know, any A-list celebrity,
like their day rate is going to be you know eight hundred thousand dollars for the day
whether they're on set or they're in the studio you're paying for their time so the second that
you disintermediate disintermediate away their time you're just paying for the value without
paying for the time and so by definition it's going to be a lot cheaper yeah so actually there
may be a world where they start to incentivize say hey actually we don't want you we want your
voice and therefore it's way cheaper to be able to do and the other thing that you'll see is you
know marvel and disney and all these other studios like they'll shoot and then they'll need to reach
like re-invite the actors into the studio in order to you know add dubbing add background audio um
you know all that jazz and like you're still paying the day rate and so what you'll have is
in the contract in the future it'll say hey we are paying you for your likeness and your image
and your time on camera but we're also paying for an exclusive license for this movie to generate
additional versions of your voice because it just allows the process to be faster. And this
obviously has to happen with permission. It should not happen without permission. You're still the
person who owns your likeness, who owns your voice. You just need to license it. And in the
same way that you license someone to use your photo, you can license someone to use a likeness
of your voice. So another thing that becomes pretty interesting is let's say that I was to
put my voice on a platform. Somebody was to come and license it. I know that it's going to go into
x movie let's say uh and then all of a sudden i don't like the script and they made me say
something that i wouldn't have said or i didn't want to say i did license it to them but there's
nuance now how do you see that stuff playing out and that may be kind of an extreme example for
movies i'm sure um it's exactly the same way as if you sign yourself up to be in a reality tv show
and they can cut it however you want in reality tv shows sometimes they'll shoot over your shoulder
you're saying one thing and then you see you know the back of my head but to take a clip of me
saying something else from a completely different thing and the person's like do you want to kill me
and i said of course not but then they switch of me saying yes right they have creative license
and you sign that away when you participate in any movie in any tv show in any whatever so it's
not that different it is true that you should always only partner with people that you actually
believe are good people and if you are partnering with people who you think are bad actors to begin
with, then that's what's going to happen. Absolutely. Talk to me about the company
itself in terms of how big is it revenue-wise, number of employees, what data or information
can you share with us? Yeah. So I started working on Speechify back in 2015. And the reason I was
excited about it is I read a bunch of academic papers about the narrow applications of deep
learning, specifically around speech synthesis. And my thesis was, we're going to have text-to-speech
be more than 10x better than it currently is like this is absolutely insane and i ended up writing
this 30-page paper about my world views and the conclusion was that i'm the person i am today
because of my experience overcoming dyslexia when i was a kid and because i listened to audiobooks
a week and i've done that for like the last 17 years 100 books per year and i was like this is
the most important piece of technology for me to invest my time in and this is why i'm on this
earth is to solve dyslexia to solve adhd um and uh i'm so lucky that audiobooks existed when i was
nine, 10, 11 years old, because if I was born 15 years earlier, 20 years earlier, I would not have
had the academic experience that I had, but not at all. I'd be a completely different person.
And yet most textbooks did not have an audio book. And so that's what we set out to do. Now
there's about a little more than a hundred people who work at Speechify. We're like 75 engineers,
30 people on the AI team. My brother, Tyler is actually blind in his left eye. He's astigmatic
in his right eye. And so he does all of his coding on a giant projector, but he started coding when
He was in third grade building triangle ballsy websites.
He taught himself assembly in fifth grade to hack video games.
He skipped three and a half years of math in high school,
three years of computer science.
And then he did math as an undergrad at Stanford.
He did his master's in AI there.
And he joined the team and leads the AI team for Speechify.
And so we both had this very big passion of making sure that reading is never
a barrier to learning for anyone, no matter what your background is.
And really above everything else,
my goal in life is to be the person that I needed most when I was young.
When I was young,
the thing I really needed was someone to do my readings for me.
And so in the beginning, it was a bunch of people like me, people with dyslexia, ADHD, low vision, autism, concussions, anxiety, second language learners.
Now that's less than, you know, 30, 20 percent.
Most users of Speechify are doctors, lawyers, accountants, people in the military, executives, people in finance who they like to listen when they drive.
They like to listen when they work out.
They like to listen when they walk.
They like to listen with their cook.
They find that they can listen two to three times faster than they can read.
They can comprehend better when they listen and read at the same time.
It's way less boring.
So the contingent of people who have used it the most happen to be people with ADHD.
What's interesting is if you have ADHD, your brain is wired slightly differently.
Turns out you're way better at listening fast.
So the percentage of people who I meet who listen at 3x speed, it's like 80% of people
with ADHD because the speed of the listening is equal to the speed at which their mind
is working and that helps them.
But if you intake the text at 200 words per minute, there's a dissonance there.
You get bored.
You act down in class.
So you walk out to, you know, pet a dog, whatever it might be.
But the second that it's equal to your brain, which already moves fast, you're absolutely
locked in.
And so we have the Chrome extension is, you know, our most popular product and then the
iPhone app.
And then we have the mobile web product and then the B2B tools that we built.
And so right now we're like 25 million people.
The goal is to go to 2.5 billion.
The thing is, audio is a fundamentally better user interface than digital screens.
And it's funny because almost every futuristic movie that you see, the predominant user interface is audio, right?
People talking to their computer and the computer will have wit and it'll be funny and it'll have a nice voice, whatever.
And so our goal is to build an auditory operating system.
We started by solving a very single player mode problem, which is how do you read documents, emails, chat, GPT, Google, physical books, and have it at Kindle and have it in an exceptional manner with super high quality voices.
all the integrations the ability to change speed to play pause to pick the word that you want to
listen to and for the most part like we've cracked that we've been the number one app in our app store
category for about four years above the new york times above the wall street journal and now a lot
of the investments that we're doing is hey users listen to so much content on speechify there will
come a time where we can recommend to you the most relevant material to you so you have all
these beautiful books behind you let's assume that you finished reading 80 of them and some of them
you stopped 25 of the way through i can start recommending to you the most relevant books for
to read next based on what you actually completed exactly like tick tock recommends videos and so
except for entertainment it's for business and knowledge purposes that to me is really really
exciting it gets even more exciting when you put it in the context of a team so now we're launching
speechify teams that knows the entire knowledge base that all your team has because it knows that
you read this macroeconomics paper about inflation as it relates to solar panels and i read this
paper about gpus from nvidia and actually we should be the two people called into the organization to
to talk about this topic.
And so that to me is just like a very exciting future.
When you think about audio,
everything we've talked about
is kind of the input to the human, right?
Audio is coming in, I'm digesting the audio.
There is the audio interface though.
We see this with Siri, we see this with Alexa,
we see this with a bunch of these different products.
Is it just going to be kind of a bimodal
or kind of two directional, I'm consuming audio,
but I'm also using audio to kind of dictate my life
and that's the dominant use case kind of moving forward?
so if you want to build an ai company and there's like 10 fastest growing generative
ai companies obviously opening eyes number one speechify is off the number like six or seven
um you've got to think where value accrues right and so there's two ways of thinking about it the
first one is does value accrue to the person who has the best model um or does it accrue
in another place and often that place actually ends up being distribution because if you have
the largest amount of distribution you get to train your models on more and more data than
everybody else um and so there were many companies hundreds of companies that were at the same level
of open ai during gpt1 gpt2 gpt3 the thing that really changed everything is chat gpt
for speechify we started with being user first and then we leaned really hard into the deep learning
part and the reason it took a little bit of time is one i was a college student and two audio is a
lot more uh heavy than text and so obviously what you're going to see is first you have this bottle
between open ai and anthropic then you'll have this bottle bottle to see who is going to win
the audio and obviously speechify and open air two leaders there and then you're gonna have this uh
fight about like text to video because that's the next thing that is the most heavy
um and so it's not enough to just have the foundational models and by the way it's important
to note that within large language models in ai they're not all the same some of us are for
textual reasoning and some are for audio and some are for video and some are for other things
so we're especially good at the voice part um but it's not just enough to have the foundational
models you need to have the end relationship with a user and so the way that speechify has decided
to build the products that we create is they sit on top of other software so if you think about
the technology stack and you think about microsoft they started with dos and then you so you start
with an operating system and then on top of that you build applications call it the app store or
the chrome browser and then you build websites on top of that and then on top of that you get to
have chrome extensions like speechify or mobile safari extensions like speechify or gmail plugins
like speechify and then we end up being the conduit for information intake we're the last
piece of technology you interface with before the information gets into your brain and then that's
the input and now we're starting to work on the output as well so the output is important because
it helps you navigate it needs to be able to deal with interruptions it needs to be able to deal
with sound in the background and by the way what's very interesting is if you're able to use speechify
technology and you're recording a conversation like this let's say i was in the middle of a
stadium, I could transcribe my voice. And then I could text the speech, my voice again, with the
same intonation, generate the voice perfectly, but drop all the background audio, drop all the
additional speakers, drop if we're in a stadium, the bouncing of the basketball, the sound, the
roaring of the crowd, whatever it might be. And so you end up with these fundamentally new primitives
for data on the internet. And that's really, really exciting. And so what ends up happening
is the user interfaces change, you're going to have a situation in the next like two, three,
four years, where your AirPods will be LTE compatible, and you can leave the house without
your iphone and that's already kind of amazing because you're not addicted to this physical
screen anymore and humans a lot of the depressions that we're seeing these days is people being
sucked into these social media platforms and you know whatever it might be tiktok instagram but if
you are able to still get all the value all the knowledge all the entertainment all the learning
without being stuck with your face inside of here but you can look up and you can walk around you
can be in nature and you can row your boat whatever it might be while listening that's a far better
human experience and you can cook and you can do it and the second that you want to boom you take
out your earphones you interface with your child and so that i think is a much more exciting feature
than where we've been headed in the last 10 years my last question for you is a little bit more fun
uh music obviously people have heard you know the con a uh renditions and many others um those were
people who basically were doing it more for entertainment purposes and kind of to show what
the technology could do do you think music artists will begin to use this technology to basically say
Hey, look, screw it. I can just, rather than go in and record myself, use my own voice and figure
out how to make, you know, chart topping hits. Of course. So we already have a bunch of artists
who work with us. Um, and it's really funny actually. So like one of the really good example
is a guy who, um, there's people who they don't want to record in the studio because the studio
sucks. They want to record on their phone. Um, so they're recording your phone, but what happens
is, you know, we'll capture the intonation, everything you have, and then we'll spit it
back out but it'll be super high quality in a lossless format because we regenerated it from
scratch so actually it starts to not matter what microphone quality you use and so you can record
in the subway as long as the vibe is right and you get the intonation that you want boom you can get
it there um and the way that the music industry works is you have a bunch of people who will write
hit songs and then artists will try it on during a studio session and they'll keep the one that
they really really like and so now what you can have is every artist they can just listen to a
a stable of songs, sang with their voice,
and then they can decide,
before they even actually go to record,
whether they like it or not.
And you're gonna have artists
who already have a really big name
be able to feature on any song you want,
as long as they give permission,
and that's really interesting.
And you obviously have artists like Grimes,
who are publicly allowing anybody to use her voice,
and just like, you know, 50% split it.
You'll see more things like this.
Again, it's not a solution for creativity,
but it is a solution for more efficient execution.
Cliff, where can we send people to find you on the internet or learn more about Speechify?
Easiest place is follow me on YouTube or Instagram, Cliff Weissman, C-L-I-F-F-W-E-I-T-Z-M-A-N.
I also write a lot on Medium.
If you want to find Speechify, search on the app store, Speechify, Speech, and then I-F-Y,
or go download the Chrome extension for Speechify.
Use it to listen to your emails.
It'll save you a lot of time.
Amazing.
We'll do it again in the future.
Great.
Anthony, great to meet you.
We'll be right back.
