Chris Almaguer Data visualization design

Case study - Tooling & on-device AI

Twenty years of music, and shuffle was the only control I had

Seventeen thousand tracks, each given a genre, a mood and a description of how it actually sounds, by a model running on the laptop itself. Nothing uploaded, nothing metered, no cost per track. The point was never automation. It was getting enough structure onto a collection that I could finally ask it something specific.

macOS app On-device AI Ollama

01 - Context

We have every song ever recorded and we listen to nine of them

I have been collecting music since Winamp, through MiniDisc, cassettes, CDs, and every hard drive since. Seventeen thousand tracks, most of them bought or ripped rather than streamed.

Streaming handed everyone the entire catalog and quietly took away the thing that made a collection yours. Recommendations go stale. Shuffle across seventeen thousand songs still serves the same forty, and repeats them. There is no way to say "the sad ones," or "the ones nobody else knows," or "surprise me, but not like that." You get a list, and you get play.

So the question was never how to tag files. It was what a person could do with their own library if the library knew something about itself.

02 - The pipeline

Five stages, one of them a model

A library is imported in stages. scan walks the files, tags reads embedded metadata, art pulls cover images, analyze measures the audio, and enrich asks a local model for the things no tag contains: genre, subgenre, mood, and free-form vibe descriptions.

Only the last stage is a model, and that is deliberate. Everything that can be measured is measured. The model is asked only for judgment, on top of evidence the earlier stages already gathered. All of it lives in an Electron shell around a local SQLite index, with Ollama serving the model on the laptop itself.

The EQ room. The header says it: shape the sound, the seance plays while the music does. Vocal cut, band sliders and the meters, running live.

03 - Why not streaming

Owning it is the point

This is the part where someone asks why I did not just use Spotify. Streaming is an extraordinary answer to "what can I play." It is a terrible answer to "what do I have." A subscription is a lease on somebody else's shelf: the catalog decides what arrives, the platform decides how it sounds, and both change without asking you.

A collection is different. These are hard files, bought and ripped across twenty years: MP3, FLAC, M4A, WAV, AIFF, OGG. Lossless plays as lossless. Nothing is squeezed through a bitrate ceiling chosen for bandwidth, and nothing quietly leaves the library because a license lapsed somewhere upstream.

The quieter loss is creative control. A platform ships one interface, one EQ and one recommendation engine for everyone, and there is no seat in that room for the listener who wants to pull the vocals out of a mix, or sit somewhere else in the hall while a record plays. This app leans the other way, toward the audiophile and the track collector. The EQ is a room you stand in. The auditorium view moves the music to whichever seat you click. And genre is not a marketing shelf but something the library knows about itself, deep enough to hold regional rap with a few thousand listeners.

own
Files, not licenses. A purchase is a file on your disk. A stream is permission, and permission gets revoked at the margins every year.
flac
Quality, not throughput. Lossless plays as ripped or bought. A stream meets a bitrate budget chosen for the network, not for the record.
eq
Your room, not theirs. One interface for everyone, or your own: band sliders, a vocal cut, a seat in the hall, palettes that survive color-vision deficiency.
disk
Permanence. A folder of files has outlived every player since Winamp. It will outlive this one too.
Lyrics, even unpublished. The panel follows playback with timed lines. When no published lyrics exist for a track, a local Whisper model transcribes the audio itself, and nothing leaves the machine.

04 - The features

What a library can do once it knows itself

Seventeen thousand lines is not something anyone reads, so the interface is charts rather than rows: the collection is shown as shape first, and every feature below sits on the same idea. Structure first, then the specific thing you actually wanted.

01
Charts, not rows. Tempo, years, tags, loudness and origins drawn as charts, with a treemap and a sankey for the deep cuts.
02
Mood shelves. Boards built from the enriched mood field, so "the angry ones" is a place you can walk to rather than a search you have to remember.
03
Ask the library. Natural language over your own files: the hidden gems, the emotional run, the regional stuff nobody else knows.
04
Lyrics, even unpublished. Timed lyrics from an open database when they exist. When they do not, a local Whisper model transcribes the track itself. Nothing leaves the machine.
05
The EQ room. Band sliders, a vocal cut, a limiter with a ceiling you can drag, and a meter bridge that reads at arm's length.
06
Every format on the shelf. MP3, FLAC, M4A, WAV, AIFF, OGG. The formats a collection actually arrives in, played as they are.
07
Playlists by hand. A wall of covers curated deliberately, renamed freely, and a queue that remembers which view filled it.

The charts

10 slides, one library
The listening dashboard: Most played and When you listen side by side over library totals.
Your library, listened to. Most played and When you listen in one view: 17,126 tracks, 2,340 artists, 1,492 albums.
The genre force map: 63 genre dots sized by track count, connected by lines.
How the library connects. A force map, not a chart: each dot is one genre, its area is the track count, and clicking a genre branches its artists.
The genre treemap: 63 rectangles, each genre sized by its track count.
How the library divides. Every rectangle is one genre and its area is how many tracks sit inside it: the shape of the collection at a glance.
The origins map: 174 cities as dots sized by track count, Los Angeles largest.
Origins. 174 cities the model tied artists to, sized by how many tracks come from there. Los Angeles leads with 2,438. An empty region may only mean nobody has been asked yet.
The sankey: nine genre blocks on the left flowing into thirteen decade blocks on the right.
Genre to decade. Nine genres flow into thirteen decades, for the 959 hours of listening the library can put a year on.
The tempo histogram: bars count tracks in four-BPM slices across 16,690 analysed tracks, peaking at 92 to 95 BPM.
Tempo. Four-BPM bins across 16,690 analysed tracks. The peak is 92 to 95 BPM and 63% of the collection sits within 10 BPM of it: the zone where a set gets built without changing pace.
The years chart: one bar per release year from 1938 to 2020, with 1996 the tallest.
Years. One bar per release year, 1938 to 2020. This is when the music was made, not when it was added: 1996 is the busiest with 1,432 tracks, and the 1990s hold nearly half the dated library.
The three-column sankey: genre on the left, decade in the middle, tempo on the right.
Genre, era, tempo. The same reading with a third column, so a genre's decades and its speeds sit in one picture.
The tags ranking: the top 24 of the model's tags drawn as bars.
Tags. The top 24 of the model's tags. A track can carry several at once, so the ranking reads as what keeps recurring, not as a split of the whole.
The loudness war chart: median LUFS by year as a line, with a quartile band behind it, over 14,709 analysed tracks.
The loudness war. Median LUFS by year with the middle half of each year's records banded behind it. A wide band is a year mastered inconsistently.

05 - The data

What the library knows, drawn

Every stage of the pipeline leaves numbers behind, and the app draws all of them. Tempo, years, tags, loudness and origins describe what the collection is made of; a graph, a treemap and a sankey take the deep cuts. Every chart states what it means and shows its key, and any of them can be saved out as an SVG.

The other half is behavioral. The app keeps its own play history, so listening becomes a dataset too: when you listen, what you replay, what you skip. Most played and Most skipped run as leaderboards that can be reset and started again, and a track can be forgotten entirely.

06 - What broke

The failures worth keeping

Free and local was the constraint from the start: nothing uploaded, nothing metered, no cost per track. Holding that line surfaced real failures, and four of them taught me more than the features did.

The model invented. A small model fit comfortably in memory and worked on anything canonical, then fell apart exactly where this library matters most, underground and regional rap. Asked about a record it had never seen, it produced a confident, plausible, wrong genre instead of an empty field. A gap is visible; a fabrication is not. The fix was not a benchmark score, it was stepping up to the smallest model whose recall did not run out on the real data.

The selector did not select. "Play me a good song" searched the library for a record called "me a good song," found nothing, and said so. Technically honest, completely useless. A request built of filler is a request to be chosen for, so it now routes to a pick instead of a lookup.

The breadcrumb had nothing to say. It held a full row across the top of every view, and in most of them it carried one crumb pointing at where you already were; in duplicate review it captioned the screen "All music." It now shows over the song table and nowhere else.

A header hid the one button that mattered. Save SVG lived inside a chart header that the five aggregate charts hide, so the button that saves a chart was missing from exactly the charts worth saving. It moved out.

Every one of these is a commit, written down the way it happened.

07 - The end game

A library that outlives its players

The end game is old fashioned. An iPod shows up as a folder, the app works out what fits, and a slice of the collection walks out the door on hardware no subscription can follow. That is the ownership argument made physical.

It is also the same problem as the rest of my work. A hospital dashboard and a record collection share a shape: more records than a person can hold, and the job is giving them a way to see the whole thing and then act on one part of it.

What I take from it

The design work here was not the interface. It was deciding what the machine is allowed to assert, and then what a person gets to do with what it asserted. Measure what can be measured, constrain the shape of what comes back, and size the model against the hardest part of the real data instead of the average part.