Step-by-step how-tos for every free tool, plus short, practical write-ups on the parts of the studio that trip people up - getting keys before you've even signed in, and knowing what to expect once you're on each page. For anything that goes wrong, see the FAQ instead.
Swara Sound doesn't talk to Suno directly - Suno's own API isn't publicly self-serve, so access goes through kie.ai, a third-party platform that resells API access to Suno's models.
Swara Sound never sees or stores this key on a server - it's saved to a hidden folder in your own Google Drive and sent straight through to kie.ai with each request.
Mureka runs its own API directly, no reseller involved.
Same as Suno: the key lives in your Drive, not on a server we control.
Suno (via kie.ai) is the simplest starting point - most consistent vocal quality, priced per song, no infrastructure to manage.
Mureka is usually cheaper per song and returns multiple takes per request, so you get a choice back for roughly the same spend.
ACE-Step on your own RunPod pod only makes sense once you're generating enough that GPU-time pricing beats per-song pricing, or you want a model you can fine-tune (see the Trained sound / LoRA feature on the Bands page). It's also the only option with real setup: deploying and keeping a pod running yourself.
All three show a live per-song cost estimate on the Create page once you've added a key, so you can compare against your own usage before committing to one.
Most AI music apps put themselves between you and the model: you pay them a markup, and your songs and prompts sit on their servers. Swara Sound doesn't do that - it's an interface, not a middleman.
Concretely: you sign up for your own account with Suno (via kie.ai), Mureka, or RunPod, generate your own API key there, and pay that provider directly at their price. Swara Sound relays the request using your key and never marks it up or stores it. Your library and settings live in your own Google Drive or your own disk, not a database we run - see Privacy & Terms for the full picture.
The trade-off is upfront setup - you need an account and a key with at least one provider before you can generate anything, which the homepage walks through in three steps.
Record or upload a raw take - just your voice, no backing track on it yet - and ACE-Step builds a brand new song locked to that recording's own melody and timing (its "cover" task). What comes back from that first step is the engine's own attempt at singing your part over a new band - good for judging the arrangement, not the final result.
A second, separate step puts your actual voice back on top: Demucs pulls the new band out from under the engine's own vocal, and your real recording gets remixed onto it and mastered. That's your voice, not a synthesized copy of it, and not the engine's re-sing - the whole point of this page over a plain text-to-song generation.
This needs the Combined RunPod image set up on Settings - one pod that runs both steps (the same account/pod Create's ACE-Step option can use too, if you deploy Combined instead of a plain generation-only image). Bills your own RunPod account, same as any other ACE-Step pod.
Band is optional - pick one to reuse a saved style prompt and mastering target across songs, or leave it for a one-off. Pick a provider first: that decides which field below it actually needs (a key for Suno/Mureka, a running pod for ACE-Step).
Suno and Mureka run in the cloud - a song is usually ready in one to a few minutes, no setup beyond a key. ACE-Step runs on your own RunPod pod, so it needs that pod actually running first; Swara Sound pings it before submitting and tells you plainly if it isn't reachable, rather than failing silently after you've clicked Generate.
A finished song saves to your Library automatically - there's no separate "save" step. If you reload the page mid-generation, a "Resume" card picks the job back up instead of losing it.
A Band is a reusable style prompt plus a mastering target (LUFS and crest) - the fix for every AI-generated song sounding like a different artist. Every song made for that band reuses both automatically.
Bands live in a hidden app-data folder in your own Google Drive, not a database Swara Sound controls - see Privacy & Terms. You need to be signed in with Google for a band to save; without that, changes only last until you close the tab.
This computer uses your browser's File System Access API to read/write a folder you pick directly - fastest, but Chromium-only. You pick the folder once and it's remembered: on later visits it reconnects by itself, or the browser asks for one "Allow" click (the sidebar chip shows ↻ Reconnect). Choose "Allow on every visit" in that prompt and it stops asking altogether.
Google Drive uses a normal, visible Drive folder instead - slower per file, but it follows you to any signed-in device and needs no reconnecting. See the FAQ's browser support question if "Choose folder" seems to do nothing.
Refresh reconciles the index against whatever's actually in the folder, so audio files added from outside Swara Sound (like the desktop app) still show up.
Batch expects an .xlsx with the exact columns in the blank template - download that instead of building a sheet from scratch. Each row is one track; leaving a row's own Duration/Language blank falls back to the default set in the side panel.
Rows run unattended in order - pending, then generating, then done or failed - and you can select just the pending ones or stop mid-run. When it's finished, Download updated sheet gives back the same file with results filled in, so re-running only touches what actually failed.
Voice training doesn't run on Swara Sound's servers or your own machine - it runs on a free Google Colab GPU, in four steps: upload your recordings to Drive, open the provided notebook in Colab, run it, and the trained voice becomes available back here once it finishes.
You only need a free Google account - no local GPU, nothing to install. How long it takes depends on which GPU Colab happens to hand you for free that session, so treat it as "leave it running," not instant.