Text-to-Speech
Voice cloning
| # | Cloning API | Sounds like100 = the person | Across voices | What that means | Audiosec | Clonecost |
|---|---|---|---|---|---|---|
| 1 | cosyvoice-v3-plusQwen⚠ | 90 | 86–95 ±5 | Closest to the original, on every voice we tried. | 10 | free |
| 2 | sonic-3.5Cartesia⚠ | 85 | 80–87 ±4 | Slightly behind on quality, but the most predictable — every voice landed in the same narrow band. | 10 | free |
| 3 | s2.1-proFish Audio⚠ | 84 | 68–85 ±9 | Matches Cartesia on a good voice and falls a long way behind on a bad one — twice the swing. | 10 | free |
| 4 | tts-rt-v1Soniox⚠ | 75 | 67–80 ±7 | Lands in the right vocal family — same age and timbre — without quite arriving at the person. | 20max | free |