I wanted the title to be as broad as possible. This may be a thread where everything related to speech synthesizers and TTS, from questions and problems to suggestions and available/newly-released voices/engines, is posted. So I have two questions:
1. I already have eSpeak-NG and RHVoice on my iPhone but can anyone list all the available third-party TTS engines? New ones may be added to this list as they're released. And does anyone know whether DecTalk will be ported to iOS?
2. Would it not be great to have an option to download third-party TTS engines from within VoiceOver settings, just like we can go to App Store to download fonts? Well, just letting me quickly go the app store and having me look for fonts or TTS engines myself doesn't actually help much, so users should only see fonts if they choose to download fonts from the App Store, or TTS engines if they choose to download TTS engines/voices.
By Enes Deniz, 24 June, 2023
Forum
iOS and iPadOS
Comments
A Request to Blind Developers
We have several blind developers on here, and I would like one of them to develop an app (s)he would actually use himself/herself rather than just try to sell. Eloquence is responsive, but certain voices like Sandy sound odd, and Apple keeps breaking and fixing things like the Community Dictionary. It does support English, but not some of the other languages supported by other TTS engines including Vocalizer. Vocalizer, in turn, has broader language support, but it has strange issues of its own, including some Apple-controlled pronunciation dictionary, not differentiating between full stops and exclamation marks, arbitrary quality degradation and weird bugs introduced by either Apple or Nuance (including weird pauses or punctuation handling). eSpeak-NG is great in certain aspects, but the developer hasn't updated it in quite a while and it might just get pulled from the App Store at any moment. It also sounds robotic, and doesn't support every included language equally well in practice. RHVoice is also in a similar situation; its developers haven't added most of the voices that are available for free on Windows, and the recently added voices require a subscription. Companies like Acapela Group keep their voices exclusive to certain apps, likely because they sign contracts with their developers. CerePlay is good, but it is not responsive enough, and it only supports a limited number of languages. Piper TTS looks promising, but it is better suited for listening to things like e-books rather than for use with VoiceOver as a lightweight TTS engine. And responsiveness is not the only problem; the voices can't spell out individual letters properly. So we now have tiny models with support for tens of languages, including Kokoro TTS, MOSS-TTS-Nano, and Supertonic-3. These models are surprisingly expressive yet still responsive, and they don't require a high-end GPU; they're designed for real-time CPU inference. The Apple Neural Engine allows even faster and more energy-efficient inference. There are several apps that bundle one model only, even if it supports everything from text import to parameter customization and voice cloning. What we need is an app that acts as a runtime for various TTS engines/models instead of only a single one, and lets us use these voices with VoiceOver or other system features. It should also support voice cloning or PDF/text import, and connecting to cloud services so that it can also be used by sighted users. This allows the developer to market the app to a larger audience rather than expecting only interested blind users to download it and then having to deal with a declining number of downloads as development and maintenance costs continue to rise. Many developers have failed in the past precisely because they tried to market their apps specifically to blind users, but we now have more mainstream alternatives to meet most of the needs that such niche apps claim to fulfill, so the gap is narrowing. The best way now is to develop your app with accessibility in mind from the ground up, but in a way that can also be used conveniently by sighted users, and looks appealing to them, instead of charging blind and visually-impaired users higher and higher prices on the grounds that your app fills a gap that no other alternative can, and that you have a smaller customer base and have to keep up with maintenance costs. Apple is greedy for sure, but devs can't justify being greedy like Apple just because we need accessible alternatives for certain tasks. So sighted content creators increasingly rely on local/on-device solutions to avoid paying for credits or recurring subscriptions for cloud services. Blind and visually-impaired users are now looking for expressive and responsive voices. And developers have started to release smaller TTS models that prioritize on-device inference but still sound decent. Such an app would therefore serve more than just blind users.
nothing beats Eloquence
Nothing beats Eloquence. I certainly will not pay for another TTS.
Nobody told you to.
Do keep in mind that not everyone uses his/her device(s) in English. Mind your own business and let those interested respond and contribute to the discussion. You don't have to say something in response to every post of mine. But if it were someone else stating that (s)he wouldn't pay for a TTS engine, I would say I wouldn't want to do that either, but Apple doesn't just let developers publish their apps on the app Store for free, and developers have to make enough revenue either through millions of global downloads, or in-app ads/purchases, or by making the app itself paid. Annoying as this may be, it is understandable from a developer's perspective.
My 2026 update
Now that we are on iOS 26 and I have an iPhone 15 along with my iPhone SE third generation as my secondary device I have a few updates to share. The bugs were VoiceOver was crashing with premium and enhanced voices seem to have been eliminated for me at least. So now on my iPhone 15 I’m using Karen premium voiceover. And my iPhone SE third generation I’m using Samantha enhanced. And I no longer have an iPad mini six generation as I traded I didn’t get my iPhone 15. So thankfully now I have been able to go back to using enhanced premiums. And then you seem to read things a bit better than they did when I was 16 so I’m happy with what’s currently available. I wouldn’t pay for another TTS engine either, but some people may have to do that if they don’t have any other options and I understand this. I agree with the comment who said that they would want the ability to delete unwanted voices from our phone. I have actually put this in as a feature request Apple before. I definitely wish I could delete the novelty voices off my devices. The only one I would want to keep a superstar. And it’s true, they really aren’t any child voices. There is kind of though because there’s no well. I know they’re kind of very jealous, but they’re kind of the only thing we have unless you want to count Apple Junior. I have used that voice before I think it’s pretty cute but we like we like people use their devices differently and people have to use what makes them happy and what worked well with them even if it’s eloquence or Grandma or whoever else people want to use. I don’t think anybody should be faulted for using a certain voice. I said I love Karen and Samantha but people have a thing I like and that’s OK.
Please Be Respectful
All,
Please be respectful when commenting on each others' posts. Saying things like "Mind your own business" in response to someone's opinion, is counterproductive and goes against everything we are trying to accomplish here.
Thank you.
Enes Deniz are your feathers ruffeled?
simply saying I won't pay for something to prevent a developer from making a product based on comments I thought it was good to get a balance. Just because your willing to pay for a product doesn't mean I am. You certainly have a right to purchase a product I have the same right to not purchase. So don't get your feathers ruffled.
@Michael Hansen
I'm also telling you to either mind your own business or just be fair and warn Dennis for his dismissive responses in reply to comments from not only me but others as well, or else I don't care about any of your warnings. Now do whatever you want. Removing my comments and terminating my acount included. Oh by the way, you better have a look at my suggestions in the iOS 27 Beta 3 Megathread and forward them to Apple. You know, they somehow survived Dennis' surveillance and oppression campaign against me. And then come back here and lecture me on respectful and constructive dialog if you still can.
PS: @Singer Girl keeps on talking about Karen and other voices she uses, in a way that's repetitive and makes no sense to me, but I don't tell her to shut up and mind her own business, because she doesn't offend me or anyone else. So I do know how to distinguish between "counterproductive" and "outright offensive". Learn your lesson, moderator.
Sorry about that
Sorry if it was repetitive. I was only saying which voices I like now because I was giving the 2026 update. Things have improved since last year so that’s the only reason why I was saying what I use now. Everybody has the right hand here to their own opinion.
so you attack the members and the moderators?
At what point is this to far?
I know it will never happen, Re: Acappella
In an attempt to bring this thread back on track...
I would love to see Acappella voices on iOS. And yes, I would absolutely pay for this. 🤔
I’ve only ever used two of those
I’ve only ever used two of the a cappella voices. I only had access to them on my victor reader stratus. The only Voice that they gave us an option for was Heather Orion. I prefer Heather Orion though. I know there’s a whole bunch more, but that’s the only voices I’ve ever had access to. I’ve never had the a cappella voices on any other devices I’ve had. I wasn’t the biggest fans of those voices, but I mean it may happen. We never know what will happen in the next couple years.
I also would not mind Ivona Voice
Alas.. Amazon will not be giving those up anytime soon. 🙁
I’ve never heard the Ivanna voice
I’ve never heard that Ivona voices. Is there a sample of them somewhere? I’ve heard about them on this website here, Apple VIS, don’t I’ve never heard the voices. I guess you could have those voices if you got like the Kindle or something.
Ivona I Believe is Voice of Alexa
I believe the Ivona voices are the same as on Alexa. 100% agree they would sound great on iPhone, at least the original ones.
correct, @Michael Hansen
Ivona Voice used to be on Android. Early Android. Like during the days of SVOX, we're talking Android 2.2 and a little later. Amazon purchased it sometime ago, and now use it for both Alexa, and VoiceView, which is their screen reader for Fire devices.
Follow this link for a short introduction and demo:
https://youtu.be/Ub6XX5ckfo8?is=QATA9r0GQJGM736n
That’s the same Voice on the milestone recorders
That’s the same Voice that the milestone digital recorders used to use. One of my friends had one of those was back in like 2007 or 2008. Is that same Voice. That’s a pretty awesome Voice. That would be pretty cool to have that for voiceover. That’s not a bad Voice at all. Thanks for the demo.
my thoughts
It’s very easy to say, nothing beats Eloquence, but I don’t think it’s really that simple.
Like a lot of blind people who grew up with robotic voices, I started with DoubleTalk. Would I go back to it now? Probably not, but I still don’t mind older-style voices.
On iOS, I usually use Fred. It’s good with English, but it doesn’t handle Turkish words well at all. When I’m going through my music library and can’t recognise a Turkish song title, I switch to one of the Turkish voices, usually Yelda. There’s another one called Cem, but I don’t like the way it sounds. That’s just personal preference, though, and that’s really the point. Everyone is going to prefer different voices.
Text-to-speech voices are probably the blind equivalent of fonts. Sighted people have favourite fonts, while blind people have favourite synthesised voices.
On Windows, I use eSpeak NG with NVDA. I tried using eSpeak on the phone as well, but I didn’t like it as much. It paused for longer on certain words than I’m used to on Windows, so I went back to Fred.
I don’t usually read large documents on my phone because I sometimes bump the screen and lose my place, but that’s more of a me problem. I’d still like to see more voice choices on iOS.
Android seems to offer far more options. Companies such as Acapela have their own apps where you can buy and download different voices. It’s almost like having a little voice shop inside the app. Apple doesn’t seem to go for that model as much.
I’m guessing you’d need a decent number of voices before it was worth putting them all into one app. I’ve heard five or more mentioned, but I’m not a developer, so I don’t know what the actual requirements are.
Would I use every voice? Probably not, but you can never really have too much choice.
The other problem with iOS is that voices are often locked inside individual apps. You can end up with the same voice installed three or four times because several apps each include their own copy. It would be much better if you could buy a voice once and use it throughout the system.
I know Apple likes sandboxing everything in the name of security, and there may be technical reasons for doing it that way. Still, it feels like there should be a better solution than every app having to bundle its own copy of the same voices.
Anyway, it’s an interesting thread, and I’ll be keeping an eye on it.
It's funny about Eloquence.
People always do this, "oh it's what you grew up with"! Sort of, not for me, I started pretty late, my intro to synths was the Echo 2. But here's why I keep using Eloquence. For most things, it works the best.
What do I mean by that? There was a subject for a post here a little while ago that said something like:
Making a hotkey for audio volume.
If you read that with Eloquence, it reads fine. This is on the Mac under 26.5.2. Switch to Samantha and read it. In case you don't feel like doing the experiment yourself, or this doesn't happen for you, I don't hear hot-key. I hear "hot". So it's literally read as:
Making a hot for audio volume.
Now, I've used the Vocalizer voices before, e.g. on Windows. So I'm just going to assume this is some kind of bug. And this is the more extreme version of things. But there are also weird pronunciation issues and all. This does go back to being used to Eloquence a bit, because of course every synth has issues. But when switching synths, I just find I run into more things where I have to stop and try and figure out what a word is or whatever.
So since I have Eloquence, I just use it. I don't think it's the best, or unbeatable, or whatever. Obviously if it went away and I had to switch synths, I'd deal, I used Vocalizer on Windows for a few years with NVDA, but I also bought Eloquence when it became available.
Anyway, my point to all of this is that IMO there's more to it than "it's like fonts for sighted people" or "it's just what you're used to". Those are both perfectly fine points, and true IMO. But I also think there's a bit more going on than a simple preference, you just happen to like it more.
The languages are nice though. I've got a Norwegian voice set up so I can try to figure out how to pronounce the names of Norwegian munnharpe tunes. That's another place where Eloquence fails, although it is better on Apple than Windows. At least on Windows, Eloquence wouldn't read things like Greek or Hebrew characters, e.g. on WIkipedia. But the Vocalizer voices would, as long as you went through character by character. This is why it doesn't make sense to say something like Eloquence is the best or whatever. At what? It really depends on what you're doing.
Fundamental elitism
I've read so many comments, here and elsewhere, where people praise eloquence. Where people start arguments over eloquence. Where people defend arguments over eloquence. Where people naysay others who do not use eloquence. I call these people.. 'fundamental elitists'. It's not just with synthesized speech either, when it comes to software and technology, you could find this all over the world. The principal can be simplified as follows:
"I use it, therefore it's the best."
The sad fact is, nothing is really.. the best.. when it comes to software and technology. There is only.. "What works for me".
Like mini, I too began using synthesized speech with eloquence. For me, it was Jaws 13 or 14, with eloquence. Running on Windows 7, no less. It wasn't too long before I started using the Vocalizer voices within JAWS, and haven't really ever looked back.
These days, I can only stand to use eloquence when I am doing anything with software code or while working within a command line environment. Otherwise, I need more clear sounding speech, such as what Vocalizer offers.
@Khomus,
You do know you can correct the pronunciation of specific words with Vocalizer voices, right?
We all have to use what works for us
We all have to use what works for us. What makes us happy. It’s gonna be different for all of us and we should just I’ll be happy and respect each other‘s choices. I don’t think there’s any best synthesizer at all because all comes down to our preferences and our use case you guys already know which ones I like.
Eloquence and the nostalgia factor
I've been keeping an eye on this thread, and what strikes mabout the Eloquence argument is that no one is actually right or wrong on any of this. Eloquence does, in fact, tend to get some pronunciations correct by default. it also tends to get some wrong, as noted. It also sounds robotic, demonstrably so, and to people who use it that's a feature, not a bug. I myself grew up on Eloquence with JAWS, but I find I don't use it much, to the point where I eventually stopped downloading it to my Windows machines. though I'm happy to have the choice to switch to it on my Mac on occasion, I don't really like using it on my phone. to me, the more human-sounding voices just sound more pleasant for long sessions, and I do a lot of long sessions. I suspect a lot of the hard lines drawn about Eloquence are about nostalgia, the same way lots of people prefer 8-bit music in games--it's nostalgia. nothing wrong with that, but that means it's purely a preference, not a hard line that others must adhere to or else. Just as an example, the VocalIzer voice I'm using appears to have a problem pausing for punctuations that aren't periods. yes, it's annoying. the fix? use Eloquence, which does a much better job of this. if I was editing a document I would switch to one of the 1Core voices or Eloquence, because where precision matters, they do the job better. But on the whole I prefer the VocalIzer voices.
That’s a really good point
That’s a really good point. I tend to prefer vocalizer voices for long reading too. I get what you’re saying. I feel like they were kind of made for that to be able to read long blocks of texts. I would love to have one of the vocalizer Voice just read me a book. That would be really fun. I know there’s a way to do that, but I haven’t figured out how to use the books app yet. I hope I will someday though. I think it’d be awesome to have one of my favorite vocalizer voices. Read me a book. I just haven’t done it yet so I wanna try it out. I grew up with eloquence too, but I switched to the vocalizer voices as soon as I knew that was an option and they were responsive enough to use in jaws so Jaws users the same voices for me as my regular phone guys. Although I really wish we had premium high variance of vocalizer voices. We don’t have those quite yet. Hopefully, someday our phones will give us enough processing power to allow for this. I’m still thankful for what we have now though. And hey, every once in a while, I kind of like using some of those older Apple voices. I didn’t get to experience those as I never had a Mac growing up. So I never had access to like Friday, Victoria and Bruce and Agas and stuff like that. I kind of like having options for that on the phone. I can’t really use those for a long time, but I don’t mind once in a while switching over to those. One of my friends had a great point. She said it was so weird that on Apple product you have to download Apple voices. She just figured like voices like that such as the Agnes Bruce speaking Victoria, I would just be preinstalled instead of you having to download them. I mean, I kind of see her point. I think that should happen to you because those files are so small. They wouldn’t take that much to pre-install them. She also said that all of the compact voice that should be preinstalled so instead of having to go into downloading regular variance of the voices we would just have to download premium enhanced voices. I think that’s a really cool feature. What do you guys think?
Re: Eloquence, pronunciations.
Of course you can reconfigure how words are pronounced. But I think there's a difference between, say, snow being pronounced like 's' followed by the word 'now', and entire parts of words like hotkey being missed completely.
Re: Best.
Agree, and I think I said this in my post, there is no such thing as "best". Not even for an application like reading books. Some people want human-sounding things, some people want crazy speed because they're apparently into plowing through things as fast as possible for some reason, it just depends on what you're doing. If I really wanted human-sounding voices, I'd probably set up a pronunciation rule, assuming you can apply it to all Vocalizer voices.
But again, my point was, it's not entirely nostalgia. There's a dude who is all into bringing back every speech synth he's ever heard of. I have a fondness for the Echo 2, demo here, not a very good one though.
https://www.youtube.com/watch?v=rlQb4BsVlmA
That is, as I think I mentioned, the synth I started with. This is the slow speed, it had two, slow and fast. I'm not clamoring for it to be brought back into modern environments so I can use it with my fancy new screen reader. I think that would be pure nostalgia, because let's face it, it's not really that great, as synthesizers go.
But it's better than eSPeak, am I right? :) Speaking of that, it, eSpeak, probably eSpeak NG, gets used in Schwung, the software that gives the Ableton Move a built-in screen reader. That's because it's small and portable and exists for the OS, some sort of Linux variant, I think. It's wonderful. So yeah. You just need different synths for different things sometimes.
You can't get locked into, this is the one true synth. Because that just ain't true. Even if you love it to bits, it's going to depend on where you're using a synth, sometimes, your favorite won't be there, and you'll have to deal. In fact, to get used to eSpeak for Schwung, I actually installed it on the Mac so I can use it occasionally for more extended periods of time.
My first computer had keynote gold
My first computer had keynote gold Voice. I honestly wouldn’t mind that coming back again. But I’m totally OK if it never does either. It is what it is. My first computer also ran windows 31 is operating system. I sure wouldn’t want that back again. But we all like what we like and it’s definitely true that there are certain times in context that she might need certain sense. I mean, I’m sticking with the ones that I stick with now but I can totally understand that you would need different ones for certain operation depending on what you’re doing. I just hope that my favorite voices will always be on the ground. If not, then I will just have to find ones to get used to. But for now I’ll stick with the ones I like. I meant the phone. Lol.
Khomus
Khomus said,
I did not quite understand this comment, but on my iPhone using the Samantha compact voice, the word "hotkey" sounds like it is supposed to thanks to the pronunciation dictionary. 🤷
Re: pronunciation.
OK. Suppose we have a word, Trinitron, which is an old TV brand. Most people have probably never heard of this. You're using speech yeah? And your synth says "trin". That's it.
If your subject is "hotkey for volume", and your synth reads "hot for volume", and we're on Applevis, it's pretty easy to pick up from context one of two things.
1. Somebody mistyped, and wrote "hot" instead of "hotkey".
2. Your synth is mispronouncing something you can fix with the dictionary.
Do you see how, if your synth is saying "trin" instead of "Trinitron", you're dealing with something else? You probably don't have context. Again, if your synth says s-word instead of sword, you can fix that pronunciation, and you can probably pick it up from context, or at least, pick up that something's wrong, and go read it character by character, so you can fix it.
If you have a totally unfamiliar word like Trinitron, and you've just heard it with a synth, it's unlikely that you'll pick it up. You'll just assume it's trin. That's the difference between a synth just mispronouncing something, saying snao instead of snow, and just quietly killing parts of words. One of them, you'll likely get from context, you'll realize it's saying snow weirdly. If it's just ignoring chunks of words you've never heard before? Probably not so much.
So yes, technically, if you know there's an issue, and you like the voice, you can just fix it. Me, I like Eloquence. And it doesn't have issues like that which I need to fix. And I think silently ignoring parts of words is different from an issue of a simple mispronunciation. And my whole point is that this is an example of a reason beyond nostalgia, which sure, is absolutely part of it, well, more what you're used to, but still. The general point is that synthesizers don't perform the same way as each other, and there may be reasons beyond "I just like it" that make you want to use one over another.
Every time Eloquence comes up people just go "oh people are used to it", "nostalgia", "that's what they grew up with", as though if those things weren't true, there would be no reason to use it.
Yes, those reasons exist for quite a few people. But there are other reasons that aren't those that also exist. Vocalizer synthesizers apparently completely ignoring parts of words is one of those reasons, a reason everybody can objectively verify for themselves.
Whether they fix it with the pronunciation dictionary or whatever isn't the point. The point is that there are real differences between how synthesizers work, and those differences can also make people prefer one over another, see also constraints like Schwung, based on Linux so using a synthesizer that is open source, free, and already built for Linux, unlike either Eloquence or the Vocalizer synths.
Not everything can even be fixed with the pronunciation dictiona
Not everything can be fixed with the pronunciation dictionary. I just discovered this one trying to get voiceover to say my last name correctly. It just can’t at all, at least not in any of the vocalizer voices. It’s kind of funny because the only synthesizer that I know for sure that can pronounce my last name correctly as eloquent. There’s a couple other voices that can too, but they’re older Apple voices. I just find that kind of funny so since I don’t really like to use those other voices as much I just deal with the fact that vocalizer can’t pronounce my last name properly and I just know what it’s supposed to say anyway and I just move on, but it is kind of funny. It’s gonna be a while. I think before we have computer voices that aren’t going to have issues pronouncing things. Cause even the AI voices and things still don’t pronounce things totally correctly. We’re just not there yet with everything being pronounced the right way and having natural inflections and more human sounding voices isn’t quite very yet either so we’ll just have to see what happens in a few years this is technology to improve.