Voicer is in free beta — credits are granted on request. Request access

Free during beta · credits granted on request

Text to speech that starts talking instantly

Fifty natural voices across nine languages, rendered on our own hardware and streamed sentence by sentence — so the first words are already playing while the rest is still being generated.

No card, no API key. Tell us what you're building and we'll top up your account.

Heart · English (US)

Grade A · 24 kHz mono

chunk 3 / 7streaming

first sound 640 ms · 412 credits · 27s of audio

af_heartbm_georgeff_siwishf_alphajf_alpha
voices
51
languages
9
output
24 kHz
to first sound
< 1s

How it works

Four steps from signup to sound

01

Ask for credits

Tell us who you are and what you want to build. We read every request by hand and top up your account — usually the same day.

02

Sign in with your email

No password to remember. We send a one-time link; clicking it drops you straight into the Studio.

03

Paste, pick a voice, generate

Choose a language and voice, tune the speed, and hit generate. Audio starts playing on the first sentence and downloads as a WAV when it's done.

04

Only pay for what renders

One credit is one character. Credits are held up front and settled against the audio actually produced — stop halfway and the rest comes straight back.

What you get

Small, fast and honest about what it costs

Streams as it renders

Your text is split on sentence boundaries and synthesised chunk by chunk. Playback begins on the first chunk instead of waiting for the whole passage.

Nine languages

English (US and UK), Spanish, French, Hindi, Italian, Portuguese, Japanese and Mandarin — each with voices tuned for it.

Credits you can read

One credit, one character. No opaque token maths — the counter in the editor is the exact price of what you typed.

Speed control

Slow a narration down to 0.5x or push an alert to 2x without changing pitch or re-recording.

Clean WAV out

24 kHz mono PCM, no watermark, no attribution requirement. Download it or wire it straight into your pipeline.

The roster

51 voices, nine languages

Grades come from the Kokoro model card and reflect how much audio each voice was trained on. A and B hold up best across long passages.

Heart

af_heart

A

Female · Flagship quality

Bella

af_bella

A

Female · Flagship quality

Nicole

af_nicole

B

Female · Very good

Aoede

af_aoede

B

Female · Very good

Kore

af_kore

B

Female · Very good

Sarah

af_sarah

B

Female · Very good

Nova

af_nova

C

Female · Good

Sky

af_sky

C

Female · Good

Alloy

af_alloy

C

Female · Good

Jessica

af_jessica

C

Female · Good

River

af_river

C

Female · Good

Michael

am_michael

B

Male · Very good

Fenrir

am_fenrir

B

Male · Very good

Puck

am_puck

B

Male · Very good

Echo

am_echo

C

Male · Good

Eric

am_eric

C

Male · Good

Liam

am_liam

C

Male · Good

Onyx

am_onyx

C

Male · Good

Adam

am_adam

D

Male · Experimental

Santa

am_santa

D

Male · Experimental

Voice previews land in the Studio once you're signed in — generating a sample costs the same credits as any other clip.

Pricing

Free while we're in beta

There is no billing in Voicer today. Credits are granted by hand so we can keep the hardware honest and talk to everyone using it. When we do start charging, you'll hear it from us first — and nothing you've already generated will be clawed back.

Open betano card required
Freewhile in beta
  • Credits granted on request, topped up on request
  • All 50 voices and 9 languages, no tiering
  • Speed control from 0.5x to 2x
  • WAV download with no watermark
  • Direct line to the person who built it

One credit, one character

Credits are held when you press generate and settled against the audio that actually rendered. Cancel mid-stream and the balance comes back.

a 15-second clip
225
a minute of narration
900
about 16 minutes
15,000

We suggest keeping clips around 15 seconds (~225 characters). Longer is allowed — it just costs proportionally more and takes longer to render.

For developers

An HTTP API

Coming soon

Bot and server access is not open yet — during beta every generation goes through a signed-in session so we can keep the hardware from being scraped. The endpoint shape below is what we're building towards, and the reference is already written.

preview only
curl https://api.strophic.in/v1/speech \
  -H "Authorization: Bearer $VOICER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "The build finished in ninety seconds.",
    "voice": "af_heart",
    "speed": 1.0,
    "format": "wav"
  }' --output alert.wav

Questions

Before you ask

Why do I have to ask for credits?

Voicer runs on a single small server we pay for ourselves. Granting credits by hand keeps that box responsive and lets us actually talk to the people using it. It usually takes less than a day.

How long can a single generation be?

There is no hard cap on voice length beyond a 5,000 character safety ceiling per request. We suggest keeping clips near 15 seconds — longer text works fine, it just costs proportionally more credits and takes longer to render.

What happens when I run out?

The Studio switches to a short form asking what you were making and what would make Voicer worth paying for. Send it and we'll usually top you up.

Can I use the audio commercially?

Yes. Kokoro-82M is Apache 2.0 and we add no restrictions of our own. The audio has no watermark and needs no attribution.

Is there an API?

Not yet. During beta everything goes through a signed-in session. The reference is published so you can see the shape it will take, and we'll open keys once the hardware can take it.

Do you keep my text?

We store the first 160 characters of each generation as a preview so you can find it in your history, plus the character count for billing. The audio itself is never written to disk on our side.

Anything else — write to hello@strophic.in.

Tell us what you're building

Every account is opened by hand. Send a couple of sentences about your project and we'll load your credits — free, for as long as the beta runs.