Anime Semantle

Guessing Shounen Because It Feels Shounen

← Back to the game

"This feels like a shounen title, so let me try another shounen title." It is a common move, and one I assumed would work. Measured, it does not: shounen titles are not particularly similar to each other.

The MAL data has five demographic labels — shounen, seinen, shoujo, josei and kids. We measured two things with them.

1. Isolation by demographic

First, the same measurement as the genre article: each title's nearest neighbor with its own franchise excluded, grouped by label.

DemographicTitlesMedianBelow 0.7
Kids60.71433%
Seinen1080.72236%
Shounen2600.72236%
Josei140.73421%
No label3730.74727%
Shoujo420.78817%

One thing already shows. Shoujo (0.788) is the easiest and shounen and seinen (both 0.722) the hardest. And the 373 titles (47%) carrying no demographic label at all are less isolated than most labelled ones.

2. Do the labels form real clusters?

Isolation alone cannot tell you whether same-label titles are close to each other. So we computed the average similarity for every pair of labels, using only single-label titles and excluding same-franchise pairs.

ShoujoJoseiShounenSeinenKids
Shoujo0.6130.5720.5220.5060.490
Josei0.5730.5940.4890.4950.483
Shounen0.5070.4710.4900.4720.442
Seinen0.4900.4820.4780.4740.455
Kids0.4800.4830.4500.456

The highlighted values are the diagonal — same-label averages. Kids-to-kids had fewer than 20 pairs and is omitted.

Shoujo and josei are real clusters

Shoujo-to-shoujo 0.613, josei-to-josei 0.594. In both cases the own-label value is the highest in its row, and the two are close to each other as well (0.572 / 0.573). Shoujo and josei genuinely form one neighborhood.

Shounen and seinen are not

shounen-shounen 0.490 < shounen-shoujo 0.507

seinen-seinen 0.474 < seinen-shoujo 0.490

In other words, a shounen title resembles a shoujo title more than it resembles another shounen title. Same for seinen. Neither has its own label as the maximum of its row.

The reason is scale. Shounen's 260 titles are a third of the pool and include battle series, sports, gag comedies and isekai all at once. "Shounen" is not a kind of story — it is a note about the magazine's readership, so it does not map onto the plot structure similarity actually reads.

Shoujo's 42 titles, by contrast, are few and genuinely alike: relationships and feelings at the center, with repeating event structures.

3. How to use this in the game

4. Limits

The matrix is not quite symmetric — josei-seinen reads 0.495 while seinen-josei reads 0.482. The similarity values themselves are symmetric, but each item stores only its top 3,000 neighbors, so some pairs appear in one list and not the other. Do not trust the third decimal; read only large gaps like 0.49 versus 0.61.

Josei (14) and kids (6) are small samples, and shoujo (42) is not large. The shounen (260) and seinen (108) conclusions rest on enough data, but the shoujo-josei cluster could shift with more.

Demographic labels do feed the structural text of the similarity calculation. That same-label pairs still only reach 0.49 means the difference in plot structure outweighs whatever the tag contributes. Method on the about page.

Further reading

A Sharper Axis Than Genre runs this measurement over 34 theme tags and Which Genres Are Hardest over 17 genres. Individual cases where tags disagree with similarity are in When Similarity Betrays the Tags. Each day's answer and its nearest neighbors accumulate in the past answers archive.