Audio & Voice
How to Use ElevenLabs to Clone Your Voice & Create Audiobooks to Sell
Complete instructional guide to recording raw voice data, generating custom narrations, and self-publishing audiobooks on ACX.
In-Depth Blueprint Outline
The Power of Voice Assets and Digital Audio Leverage Step 1: Acoustic Room Treatment & Echo Isolation Step 2: Studio Hardware Stack (Microphones & Interfaces) Step 3: Recording & Vocal Conditioning (Audacity Configuration) Step 4: Setting Up ElevenLabs Professional Voice Cloning Step 5: Publishing & Monetizing on Amazon ACX (Audible)The Power of Voice Assets and Digital Audio Leverage
In the digital economy, **voice is the new leverage**. Millions of commuters, entrepreneurs, and busy professionals consume audiobooks, podcasts, and video voiceovers daily. Traditional voiceover work is linear: you spend hours sitting in a recording booth, speaking into a microphone, and editing out breathing offsets. If you make a mistake, you must re-record.
By leveraging **ElevenLabs Professional Voice Cloning (PVC)**, you decouple your income from your active hours. You record your voice once to train a neural network, and from that moment on, you can generate complete, high-fidelity audiobooks, video scripts, and podcasts in minutes by simply pasting text. This guide outlines the exact, step-by-step masterclass blueprint to configure your recording space, equip your studio with premium hardware, train your AI voice clone, and publish audiobooks on Amazon ACX to capture compounding royalty distributions.
Step 1: Acoustic Room Treatment & Echo Isolation
The single most important factor for high-fidelity voice cloning is a **clean, dry vocal recording**. AI models are extremely sensitive: if your recording has background hum, computer fan noise, or room reverb (echoes), the cloned voice will sound metallic, hollow, and unnatural.
Before recording a single word, optimize your space with these acoustic treatments:
- Minimize Hard Surfaces: Hard walls, wooden desks, and glass windows reflect sound waves, creating reverb. Close your curtains, lay down a thick rug under your desk, and hang soft blankets or heavy towels nearby.
- Acoustic Foam Panels: Place 12-pack acoustic foam panels ($25–$35 on Amazon) on the walls directly in front of and behind your microphone position to absorb mid-to-high frequency reflections.
- Eliminate Ambient Noise: Turn off air conditioning units, shut your doors, and record during quiet hours (such as late at night or early in the morning).
Step 2: Studio Hardware Stack (Microphones & Interfaces)
To train a professional-grade AI clone, your recording microphone must capture your vocal warmth and full dynamic range. Built-in laptop microphones or cheap headset mics will fail. You need a dedicated studio condenser or dynamic microphone setup.
Here are the recommended physical audio gear products to build your authority recording stack, complete with current pricing and target specs:
Shure SM7B Dynamic Vocal Microphone
The industry standard dynamic mic. Captures warm, natural speech while rejecting off-axis background noise. Requires an XLR connection and 60dB+ interface gain.
View on Amazon
Audio-Technica AT2020USB-X USB Condenser Mic
Plug-and-play simplicity for voice cloning. Captures high-resolution 24-bit/96kHz digital audio directly with zero latency monitoring and capacitive mute.
View on Amazon
Focusrite Scarlett Solo (4th Generation)
Features a studio-grade preamp with a massive 120dB dynamic range. Connects professional dynamic or condenser XLR mics to your PC with zero latency.
View on Amazon
Microphone Isolation Shield & Pop Filter Set
Folding isolation shield lined with high-density acoustic foam. Mounts behind the mic to absorb sound reflections and prevent room echo in untreated spaces.
View on AmazonIf you are on a budget, choose the **Audio-Technica AT2020USB-X**. It plugs directly into your computer's USB port, has an integrated mute button, and records at 24-bit/96kHz, saving you from needing an external audio interface.
If you want the absolute best studio broadcast sound, invest in the **Shure SM7B** paired with the **Focusrite Scarlett Solo**. The dynamic capsule of the SM7B is legendary for rejecting ambient room noise, making it the perfect tool for voice actors recording in untreated rooms.
Step 3: Recording & Vocal Conditioning (Audacity Configuration)
To train ElevenLabs Professional Voice Cloning, you need to upload a minimum of **30 minutes** (ideally 45–60 minutes) of continuous, clean vocal audio. We will use the free, open-source software **Audacity** to record and format our audio files.
1. Audacity Recording Settings
Configure Audacity with the following standard recording options:
- Project Rate (Sample Rate): Set to 44,100 Hz or 48,000 Hz.
- Recording Channels: 1 (Mono) Recording Channel (audiobooks must be submitted in mono).
- Audio Format: 24-bit PCM or 32-bit Float.
2. Recording Technique
Sit 6 inches away from the microphone. Speak in your natural, conversational voice. Maintain consistent pitch, pacing, and energy. Do not whisper or shout. If you make a mistake, pause for 2 seconds to make editing easy, and re-read the sentence. Read a diverse set of texts—such as news articles, book pages, or script drafts—to train the model on a wide range of words and pronunciations.
3. Noise Cleaning Process
Once your recording is finished, apply Audacity's noise cleanup pipeline:
- Select a 3-second segment of absolute silence where you didn't speak. Go to Effect > Noise Removal and Repair > Noise Reduction, and click Get Noise Profile.
- Select the entire audio track (Ctrl+A), open Noise Reduction again, set Noise Reduction to 12dB, Sensitivity to 6.00, and Frequency Smoothing to 3, then click OK. This eliminates steady background computer fan hiss.
- Apply **Limiting**: Go to Effect > Volume and Compression > Limiter. Set Limit to -3.0dB to ensure your audio peaks never distort.
Step 4: Setting Up ElevenLabs Professional Voice Cloning
ElevenLabs offers two types of voice cloning: **Instant Voice Cloning** (using a 1-minute clip, best for quick testing) and **Professional Voice Cloning (PVC)**. PVC trains a custom deep-learning model on your voice, mapping your unique accents, breathing patterns, and cadence. This is the only type suitable for long-form audiobook generation.
Follow these steps to train your professional voice model:
- Go to the **ElevenLabs Pricing Plan** and subscribe to the **Creator Plan** ($22/mo, which is frequently discounted to $11 for the first month). PVC is only available on Creator plans and above.
- Navigate to the **Voice Lab** and click **Add Generative or Cloned Voice**. Choose **Professional Voice Cloning**.
- Upload your cleaned Audacity audio files (ensure total length is at least 30 minutes, formatted as high-quality WAV or MP3 files).
- Provide consent by reading the dynamic verification paragraph displayed on the screen. This security step ensures you can only clone your own voice.
- Click **Train Model**. The training queue takes approximately 2 to 4 hours to complete. Once finished, your custom PVC clone will appear in your voice library.
Adjusting Voice Settings for High-Retention Output
Inside the Synthesis tab, use the following settings for narrating audiobooks:
- Stability (Set to 65%): Higher stability keeps the tone consistent across long chapters, while lower values add more emotional expressiveness.
- Clarity / Similarity (Set to 80%): Maximizes matching to your actual voice quality.
- Style Exaggeration (Set to 10%): Keep this low to prevent the model from over-exaggerating emotions during simple narrative sentences.
Step 5: Publishing & Monetizing on Amazon ACX (Audible)
With your ElevenLabs cloned voice ready, generate your audiobook chapters by pasting your book text into the project editor. Download each chapter as a separate MP3 file. To list your book on Amazon, Audible, and Apple Books, you must submit it through **ACX (Audiobook Creation Exchange)**.
1. Amazon ACX Formatting Standards
To pass ACX quality control, your uploaded files must meet these strict criteria:
- Bit Rate: Constant Bit Rate (CBR) of 192kbps or higher.
- Peak Amplitude: Maximum peak levels must not exceed -3.0dB.
- Noise Floor: Noise floor must be -60dB RMS or lower (this is why acoustic treatment and noise reduction are critical).
- Format: All files must be MP3, containing exactly 1 to 5 seconds of silence at the beginning and end of each chapter.
2. Royalty Models & Profit Maximization
ACX offers two distribution routes:
- Exclusive Distribution (40% Royalty): You distribute only through Amazon/Audible/Apple. You earn a massive 40% royalty share on every sale. This is highly recommended for beginners because Amazon dominates over 80% of the audiobook market.
- Non-Exclusive Distribution (25% Royalty): You earn a lower 25% royalty, but you are free to sell the audiobook on your own website, Spotify, or Google Play.
By pairing the **Amazon Associates Book Method** with **ACX voice assets**, you can promote your own audiobook on your website, earning a 40% audiobook sale royalty plus an Amazon affiliate commission on any other items the buyer adds to their cart!
Atomic Habits
Establish daily, high-output production loops. James Clear provides an outstanding framework to build powerful daily systems and achieve massive compounding results.
View on Amazon