Adding a “Listen” button is easy. Making it cheap and boring to run at archive scale is the hard part.
Ski Area Management (SAM) is a trade publication for mountain resort professionals. Their readership consists of operators, managers, and staff who spend their days between lifts, in trucks, and on the hill. For this audience, reading a long-form feature at a desk is rarely an option. To meet their readers where they are, SAM wanted to introduce audio narration.
They did not want a pilot program on ten recent posts. They wanted a “Listen” button on every article, spanning their entire back catalog on www.saminfo.com (you can see it live on pieces like their Alterra CEO announcement or their Vail Resorts earnings report).
In publisher media, two things typically kill the idea of full-archive audio. The first is the invoice: a per-character or per-minute bill from a commercial voice vendor that scales linearly with the size of the archive and the frequency of publication. The second is the infrastructure drag: thousands of audio files that quietly eat the web server’s storage until the disk fills, alongside rendering processes that consume CPU cycles until the site slows down.
For SAM, we needed to deliver audio across roughly 3,900 articles without creating a permanent operational burden or an unbounded monthly expense.
The compute and storage drag
When you introduce text-to-speech generation into a standard WordPress environment, you are asking a system optimized for delivering text and images to become a media production facility.
Initially, rendering the narration for the SAM archive meant processing those 3,900 articles on the existing 2-vCPU web server—the same server responsible for handling reader traffic and search queries. During business hours, this setup could only drain the audio generation queue at a rate of one article every three to twelve minutes. It was functional, but it meant the web server was spending valuable compute cycles reading text instead of serving pages to visitors.
The storage footprint was equally heavy. The narration MP3s totaled 11.3 GiB of data. In a single deployment, these audio files consumed roughly 39 percent of the entire media uploads folder. Left unchecked, audio generation would force a continuous cycle of server upgrades simply to store and generate MP3s.
We needed to move the compute off the web server and move the files out of the local storage, without breaking thousands of embedded audio players and without initiating a risky database rewrite.
Decoupling the generation
To avoid the recurring costs of a commercial text-to-speech API, we built the Listen feature into SAM’s WordPress infrastructure using self-hosted Kokoro text-to-speech paired with ffmpeg. This gave us complete control over the voice generation and the resulting media files.
We removed the rendering workload from the primary web server by standing up a separate worker. This dedicated worker operates on a pull model, fetching generation jobs over an authenticated WordPress REST API connection. The web server simply signals that an article needs audio, and the worker handles the heavy lifting of processing the text, generating the audio, and returning the finished file. The web server goes back to serving readers, and the audio generation queue drains without competing for resources.
Taking ownership of the voice generation also meant we had to take ownership of the pronunciation. SAM’s editors caught a problem we would have been slower to hear ourselves. Trade copy is dense with figures, and the model was reading “64.7 million” as “64, pause, 7 million.” In a publication where the numbers are the story, that is not a cosmetic issue.
Because we controlled the generation pipeline, we could adjust how the voice parsed and read numbers and decimals. To deploy the fix, we relied on a content signature system we had built into the integration. We identified every previously narrated article that contained the flawed number formatting and regenerated only those specific audio files. By fingerprinting the content, an article only ever re-renders when its actual words change, saving compute time and preventing unnecessary churn.
Alongside the MP3, the generation worker produces a sentence-level transcript. The transcript is synchronized with the audio, so the on-page player highlights each sentence as it is read. Readers who want to skim and listen at the same time can.
Offloading the storage
With generation off the box, the files were still on it.
We offloaded the entire audio library to Cloudflare R2, served behind a custom domain (audio.saminfo.com). Cloudflare R2 provides object storage with a distinct advantage for media delivery: zero egress fees.
To make the architecture resilient, we treat the audio files as immutable. Every time an article is updated and its audio re-rendered, the system generates a new file with a new name rather than overwriting the old one. This immutability means we can cache the audio files aggressively at the edge for a year. It also natively supports range requests, which is critical for audio delivery; a reader can seek to the middle of an article, and the browser will only download the bytes it needs. And when a render comes out wrong, the old file is still sitting there serving listeners while the new one is built. Nothing is ever overwritten in place, so nothing is ever briefly broken.
Migrating 3,900 media files off a WordPress server usually involves a massive database rewrite to update attachment URLs. We bypassed that entirely. We designed the audio player to select the delivery URL dynamically when the page renders. The switch between local storage and R2 delivery is controlled by a simple flag in the codebase, not a permanent migration in the database.
Before we flipped the flag, we checksummed every local file against its copy in R2. All 3,867 matched.
When we switched the playback to the R2 bucket, we left the local copies on the web server disk. If anything misbehaved, flipping the flag back would have taken seconds. The local copies stay for seven days after cutover, then go.
The result at scale
For SAM’s marketing and publishing teams, audio on the entire archive is a tangible feature for their readership. It meets operators where they are, and it creates a new, site-wide inventory line for advertisers that doesn’t rely on manually recording a handful of sponsored posts.
For operations, the infrastructure is quiet. The web server holds no MP3s and renders no audio. It serves pages. The cutover was boring, equipped with a kill switch and a rollback path at every step.
For ownership and finance: the complete audio library costs about three cents a month to store.
If a publisher chooses a commercial voice vendor to narrate an archive of this size (roughly 22 million characters), the initial run alone typically costs between $350 and $2,200 at list price, depending on the model tier (using publicly available pricing from providers like ElevenLabs, Google, Amazon Polly, and Azure). That cost repeats for every re-render and every new article. Self-hosting Kokoro eliminates the per-character fee.
On the storage side, 11.3 GiB on Cloudflare R2 runs about three cents a month, primarily because the first 10 GB are free. More importantly, R2 does not charge for egress. If we assume an illustrative volume of 100,000 full plays a month, a standard object storage provider like Amazon S3 would charge roughly $19 a month in egress fees alone. Keeping the files on the web server by upgrading a standard block storage volume would cost about $10 a month, while failing to solve the CPU and bandwidth problems.
Even if the archive triples in size, that R2 storage cost will remain under fifty cents a month. There are no bandwidth penalties for high listener engagement, and there is no monthly vendor invoice for character counts. Every article on the site can now be heard, and the infrastructure to support it is cheap, fast, and entirely under the publisher’s control.