About

Movie Madness Map

What this map is, how it's built, and how to read it.

A personal project by Steven Fazzio, not affiliated with Movie Madness · Data as of June 2026

Start here #

Movie Madness, in Portland, Oregon, is one of the largest video rental stores left on Earth. It's a nonprofit now, run by the Hollywood Theatre, and its shelves hold close to 99,000 discs and tapes. I'm a fan, not an employee or a spokesperson. I just love what they're doing, and I wanted to see the whole collection at once.

This map is my attempt at that. Every dot is one film (about 82,000 of them, collapsed from all those separate discs and tapes), and films that are about similar things sit near each other. Wander it the way you'd wander the aisles: zoom into a corner, see what's gathered there, and click a dot to open that film on the store's own catalog. If something catches your eye, go rent it. The store runs on memberships and rentals, and it ships nationwide.

How to read it #

The one idea that makes the map make sense: position means similarity. Two films end up close together when their descriptions use similar language, so you'll find pockets of slasher horror, Hong Kong action, French New Wave, and concert films, each in its own part of the map.

A few things the map is not. It isn't a ranking; a dot in the middle isn't better or more important than one near the edge. Distances are rough: a short hop genuinely means "pretty similar," but you shouldn't read the exact gap between two far-apart regions as a precise measurement. And it's a snapshot. The store's inventory changes constantly, so this is roughly what was on the shelves in June 2026, when I pulled the data.

What's on the map #

The collection comes straight from Movie Madness's public online catalog. That gives me each film's title, year, the formats it's stocked on (DVD, Blu-ray, 4K, VHS), its MPAA rating, and which shelf section it lives in. Many films also carry a written synopsis.

One dot is one film, not one disc. The same movie on DVD, Blu-ray, and VHS becomes a single dot that remembers all three formats. TV shows are split by season, since each season has its own shelf life, and the occasional book in the catalog gets a dot too.

Where the catalog has gaps, I fill them in from The Movie Database (TMDB), a community-run movie encyclopedia: posters, directors and cast, genres, and a synopsis for the roughly one film in three that the store didn't describe. When both have a synopsis, the store's own wording wins.

How a film finds its place #

To place a film, a computer first has to turn its words into something it can compare. The tool for that is an embedding model: you feed it text, and it hands back a long list of numbers that captures the gist of what that text is about. Two films with similar descriptions get similar lists of numbers, and "similar list of numbers" is something a computer can measure exactly.

Each film's numbers come from its title, year, and synopsis, plus the name of the shelf it sits on in the store. That last part matters more than it sounds. Movie Madness doesn't shelve alphabetically; it shelves by hand, into sections like "Hollywood Directors," "Leviathans & Behemoths," and dozens of national cinemas. Folding those section names into the mix lets the store's own taste help shape the map, which is why the neighborhoods feel like Movie Madness and not like a generic genre chart.

Those lists of numbers live in a space with far more than two dimensions, so a final step flattens them down to the flat layout you can actually look at, keeping near things near.

How the regions get their names #

The labeled regions floating over the map ("Slasher & Splatter," "Italian Cinema," and so on) aren't shelf names lifted from the store. They're generated. Once the films are laid out, a clustering step finds the dense clumps, and then an AI language model reads a sample of films from each clump and writes a short name for it. The names shift as you zoom: a broad area up close resolves into several finer neighborhoods as you go in.

Some parts of the map don't get a name at all. Where the films are too spread out or too mixed for a clean label, the region is left blank rather than forced into a category that doesn't fit. Those blanks are honest: they mark places the map couldn't summarize, not gaps in the collection.

How AI is used (and how it isn't) #

I'll be straight about where AI is and isn't involved here, because I know plenty of people are worn out by it being stuffed into everything.

The most important thing first: no AI wrote any of the plot descriptions. Every synopsis you read is either the store's own catalog text or a description from TMDB, written by people. I made that a hard rule. The obscure tail of an 82,000-film collection is exactly where a confidently made-up plot would slip by unnoticed, and a movie map full of plausible-sounding fiction would be worse than useless.

AI does two jobs on this map, and both are about organizing what already exists, not inventing anything. First, an embedding model reads each film's real description and turns it into numbers so similar films can be grouped. Second, an AI language model writes the names for the regions. Neither one is judging whether a film is good or deciding what it "really" is; they're sorting and labeling text that was already there.

And the most interesting structure on this map isn't the AI's doing. The reason there's a coherent French cinema corner, or a whole wall of slashers, is that the staff of Movie Madness spent years shelving those films that way by hand. The map mostly just makes decades of human curation visible from above.

The machine parts aren't perfect, of course. The region names are a model's best guess from a handful of examples, and some are off. The grouping can misread a thin or misleading description and drop a film somewhere strange. And I'll admit there's some irony in using closed, commercial AI models to celebrate a place that is about as physical, local, and human as it gets. I reached for them because they did the job well, but they're a tool, not the point of the thing.

Finding your way around #

What the colors mean #

By default the dots are colored by region, matching the named neighborhoods. The dropdown at the lower-left switches the coloring to any of these:

ColoringWhat it shows
Clusters (default)The map's named regions, each in its own color. This is the coloring that matches the floating labels.
Release yearA light-to-dark scale running from the oldest films to the newest.
GenreThe store's own broad genre (Horror, Foreign, Comedy, and so on). The fifteen most common each get a color; everything else is grouped as "Other." A film in several genres is colored by its rarest one, so big buckets like "Foreign" don't swallow the whole map.
Format (best available)The highest-quality format the store stocks a film on, from 4K down through Blu-ray, DVD, and VHS.
MPAA ratingG through NC-17, plus X and "N/R" for films that are unrated or not on record.
Editions in stockHow many separate editions the store carries, from one up to ten or more.
TMDB popularityHow much attention a film currently gets on TMDB, on a log scale so the few blockbusters don't flatten everything else.
RuntimeLength in minutes.

Films missing a value (no known year, say, or no match on TMDB) are shown at the neutral end of the scale, and they drop out of the matching filter the moment you adjust its slider.

Limitations #

Every map is a particular way of looking, and this one has clear blind spots. As with any map, the map is not the territory: it's a projection, shaped by what I chose to measure and how I chose to draw it. To paraphrase the statistician George Box, all models are wrong, but some are useful. I think this one is useful, but it helps to know where it falls short.

The project is called "Movie Madness Map", but I think of that as shorthand for "A Movie Madness Map", not "The Movie Madness Map". Part of it is just that I don't work there. But mostly it's that someone else could take the same collection, make a different set of careful choices, and end up with something that looks nothing like this. There are a lot of ways to arrange tens of thousands of films, and this is the one my choices happened to produce.

More concretely, a few specific things to keep in mind:

It's only as good as the descriptions. A film with a thin or misleading synopsis can land in the wrong neighborhood. If you spot a movie that's clearly out of place, that's usually why.

The match to TMDB isn't perfect. Linking 82,000 catalog entries to an outside database by title and year goes wrong sometimes, especially for obscure films, foreign titles, and movies with common names. A bad match means a wrong poster or the wrong cast on the card.

The region names are approximate. They're a useful guide to what's nearby, not an authoritative taxonomy, and the blank regions are simply areas the map honestly couldn't name.

Flattening loses detail. Squashing a many-dimensioned layout down to the two dimensions on your screen always distorts something. Trust that near means similar; don't over-read the exact distance between two distant regions.

What it's good for, despite all that, is wandering. It's a way to get lost in a great collection, stumble onto a section you didn't know existed, and walk out (virtually) with three movies you'd never have thought to search for.

Credits and thanks #

First and above all, thank you to Movie Madness and the Hollywood Theatre for keeping a collection like this alive and open to the public, and for putting the whole catalog online where a curious fan could find it. If this map sends even a few people their way, I'll be happy. Go become a member.

Posters, cast, and the gap-filling descriptions come from TMDB and its community of contributors. This product uses the TMDB API but is not endorsed or certified by TMDB.

The map is built with open-source tools from the Tutte Institute for Mathematics and Computing: UMAP for the layout, Toponymy for finding and labeling the regions, and DataMapPlot for the interactive map itself. The embeddings are from Cohere, and the region names are written by Anthropic's Claude.

I'm Steven Fazzio, a data scientist who makes these maps for fun. I'm not affiliated with Movie Madness in any way; I'm just a fan with a soft spot for enormous collections. The full source code is on GitHub, and if you find a film that's clearly misplaced or mislabeled, I'd genuinely like to hear about it. I've made maps like this one before, of a famous catalog of integer sequences, the most-starred projects on GitHub, and the most-liked datasets on Hugging Face.

Technical details

The short version of the pipeline, for anyone curious about the machinery. The full source, the exact prompts, and the catalog quirks I ran into along the way are all in the repository.

Catalog sourceMovie Madness public WordPress REST API (~98,834 catalog entries), fetched politely and cached
Films on the map82,256 (one per title and year, after collapsing formats)
EnrichmentTMDB: posters, cast and crew, genres, popularity, runtime, and synopsis gap-fill
Embedding modelCohere embed-v4.0, 1024 dimensions, input type "clustering"
LayoutUMAP (cosine distance, n_neighbors 15, min_dist 0.05, random_state 42)
Regions and namesToponymy for hierarchical clustering, with Claude Haiku 4.5 writing the region labels
RenderingDataMapPlot, plus a few hand-written touches (filter panel, mobile card, on-zoom labels)
CodeA sequence of plain Python scripts, eight stages, fetch through render