Skip to the content.

Where the connections started to mean something

This is the project where Sam Castillo found his way into neural networks and large language models. It did not start as a model. It started as a drawing of movies.

On an IMDb title page there is a short list: people who liked this also liked about twelve other films. Sam took a Kaggle table of about 5,000 movies, followed those lists, and drew a line whenever one film recommended another. A movie stopped being only a row of year, budget, and score. It became a point with neighbors.

Keywords make the same kind of picture from the other direction. Two films can sit close because they share plot words, even when neither page names the other. Recommendations and keywords are both ways of saying that a film is partly made of the films around it.

That is the click. A neural network and a language model also keep meaning in the links: a word, a sentence, or a film is a position among other positions, not a sealed box. The later math is heavier. The idea was already here, in a homemade map of who points to whom.

The drawings

These are the network files already in the repository. The PNG is embedded as it was saved. The two TIFF drawings live in larger networks.7z (big_network.tiff, 5330×3740, and Rplot03.tiff, 10000×5000). Browsers do not show TIFF, so each one is also here as a trimmed PNG and WebP. The archive is still the original.

Colored network of films linked by IMDb recommendations, from Network_genre_cluster.png.
Network_genre_cluster.png. Films are the points and IMDb recommendations are the lines. Color follows the first genre on the title, the same coding as the R notes. Labels on the drawing include titles such as Jane Got a Gun, Alice Through the Looking Glass, and Hotel Transylvania 2. Open the PNG.
Grayscale movie-network drawing converted from big_network.tiff.
big_network.tiff, drawn large and kept in grayscale: white page, gray links, black points. The file above is a trimmed copy at 1600 pixels wide so the page can show it. PNG, WebP, original in larger networks.7z.
Grayscale network plot converted from Rplot03.tiff.
Rplot03.tiff, an R plot of the same kind of network, 10000 by 5000 pixels, also grayscale. The page shows a trimmed copy. The TIFF itself stays in the archive. PNG, WebP, original in larger networks.7z.

How a recommendation becomes a line

Each IMDb title page carries a recommendation panel. Under a film such as The Shawshank Redemption, the section “People who liked this also liked” offers a window of about twelve other movies. This project turns that panel into a network. It starts from the IMDb 5000 Movie Dataset on Kaggle and, for each title kept in the working table, scrapes those recommended links. A movie becomes a node, and a recommendation becomes a chance at an edge.

The Kaggle extract is full of gaps. Gross, budget, and several other columns are blank often enough that a complete-case model would discard a large share of the catalog. The plot needs a set of titles and the links between them, so rows with missing values are dropped and the table is cut down to gross, genres, title, country, the IMDb URL, budget, release year, score, and content rating. Each film keeps the first genre in its list, which gives the node a single color, and the title text is cleaned before it is used as a label.

Gross and budget are restated in 2016 dollars. A consumer-price index is joined on release year, and each film’s dollars are scaled by the ratio of the 2016 index to the index for its own year. The network is limited to films made in the United States, so a year slice compares titles from one production country.

The scraper reads a movie page, keeps the lines that contain rec_item, and pulls the IMDb title constant out of each recommendation slot. Those constants are rebuilt as title URLs. For a chosen slice, an adjacency matrix stores 1 when film i recommends film j and 0 otherwise. ggnet2 draws that matrix as an undirected network. Color comes from the simplified genre and RColorBrewer’s Set1 palette, which is why the plot holds at most nine genres: Action, Adventure, Animation, Comedy, Crime, Drama, Fantasy, Horror, and Biography. Labeling and node size are arguments, so a crowded slice can drop the names and shrink the points.

The slices show how much of that recommendation structure sits inside a window of years. Releases from 2016 onward leave a small U.S. graph, about forty-five films, with titles still readable. The 1970s graph is sparse because many of the twelve recommendations for each film fall outside the decade, so the in-slice matrix is thin. Films from before 1975 are thinner still, and the edges that remain tend to join movies released near one another, sequels among them. A 2006–2007 slice is too dense for labels. With the names removed, the layout shows whether genre, along with year, pulls the nodes into groups. The notes title that crowded slice “Movies from 2010–2014”; the code keeps U.S. films with a release year after 2005 and before 2008. The drawn slices are saved in IMDB_network_final.pdf.

Read the notes (PDF) Source Original TIFFs (.7z)