Skip to the content.

Films and actors as keyword vectors

The table in this repository is the IMDB 5000 extract: 5,043 films, each with a plot_keywords field of tags separated by a pipe. Splitting those fields yields a vocabulary V of 8,086 distinct tags. A film becomes the binary vector x in {0, 1}V whose k-th coordinate is 1 exactly when tag k is on that film, and 0 otherwise. The only cast columns used are the three billed leads, actor_1_name, actor_2_name, and actor_3_name.

An actor profile is the sum of those binary vectors over the films in which the actor is one of the three leads. Coordinate k is then a count: how many of those films carry tag k. It is not binarized again. Actors with ten or fewer credits in those columns are dropped, so a single film cannot define the profile. The actor list in the Shiny app is that restricted set.

Cosine similarity of two nonnegative vectors x and y is cos θ = (x · y) / (‖x‖₂ ‖y‖₂), the cosine of the angle between them. Multiplying either vector by a positive constant leaves the value unchanged, so two actors with the same mix of tags and different career lengths have cosine 1. On binary film vectors the dot product is the number of shared tags and each Euclidean norm is the square root of the length of that film’s tag list, so the formula is the Ochiai coefficient |Kx ∩ Ky| / √(|Kx| |Ky|). A short list contained in a long one is not forced toward zero.

Euclidean distance is d₂(x, y) = ‖x − y‖₂ = √(Σk (xk − yk)²). It uses magnitude. For binary film vectors each disagreeing coordinate contributes 1 inside the sum, so d₂(x, y) = √|Kx Δ Ky|, the square root of the number of tags that appear on exactly one of the two films. A long keyword list sits far from a short one even when every tag on the short list is shared. On actor count vectors the same formula penalizes a gap in how often a tag is used, not only in whether it is used. Cosine compares the direction of the keyword mix. Euclidean distance compares the raw count profile.

A query adds two vectors and searches the rest of the collection. For films A and B the sum s = a + b has coordinates in {0, 1, 2}: a tag present on both films counts twice. The movie panel returns the other film z that maximizes cos θ(s, z), and it prints that cosine together with the angle θ = arccos(cos θ), once as a fraction of π and once in degrees. In the project notes, Frozen + The Expendables is nearest to The Chronicles of Narnia: The Lion, the Witch and the Wardrobe, with cosine about 0.283, an angle of about 0.409π, roughly 74°. The dot product that produces this ranking gives weight 2 to a tag shared with both queries and weight 1 to a tag shared with only one. The actor panel forms the same kind of sum from two count vectors. The nearest actor under cosine is the one that maximizes cos θ(s, z). The nearest actor under Euclidean distance is the one that minimizes ‖s − z‖₂. The explorer is the movie and actor addition app. The vectors and both rankings are built in server.R and ui.R.

How a query is scored

Open the Shiny app Source