"This feels like a shounen title, so let me try another shounen title." It is a common move, and one I assumed would work. Measured, it does not: shounen titles are not particularly similar to each other.
The MAL data has five demographic labels — shounen, seinen, shoujo, josei and kids. We measured two things with them.
1. Isolation by demographic
First, the same measurement as the genre article: each title's nearest neighbor with its own franchise excluded, grouped by label.
| Demographic | Titles | Median | Below 0.7 |
|---|---|---|---|
| Kids | 6 | 0.714 | 33% |
| Seinen | 108 | 0.722 | 36% |
| Shounen | 260 | 0.722 | 36% |
| Josei | 14 | 0.734 | 21% |
| No label | 373 | 0.747 | 27% |
| Shoujo | 42 | 0.788 | 17% |
One thing already shows. Shoujo (0.788) is the easiest and shounen and seinen (both 0.722) the hardest. And the 373 titles (47%) carrying no demographic label at all are less isolated than most labelled ones.
2. Do the labels form real clusters?
Isolation alone cannot tell you whether same-label titles are close to each other. So we computed the average similarity for every pair of labels, using only single-label titles and excluding same-franchise pairs.
| Shoujo | Josei | Shounen | Seinen | Kids | |
|---|---|---|---|---|---|
| Shoujo | 0.613 | 0.572 | 0.522 | 0.506 | 0.490 |
| Josei | 0.573 | 0.594 | 0.489 | 0.495 | 0.483 |
| Shounen | 0.507 | 0.471 | 0.490 | 0.472 | 0.442 |
| Seinen | 0.490 | 0.482 | 0.478 | 0.474 | 0.455 |
| Kids | 0.480 | 0.483 | 0.450 | 0.456 | — |
The highlighted values are the diagonal — same-label averages. Kids-to-kids had fewer than 20 pairs and is omitted.
Shoujo and josei are real clusters
Shoujo-to-shoujo 0.613, josei-to-josei 0.594. In both cases the own-label value is the highest in its row, and the two are close to each other as well (0.572 / 0.573). Shoujo and josei genuinely form one neighborhood.
Shounen and seinen are not
shounen-shounen 0.490 < shounen-shoujo 0.507
seinen-seinen 0.474 < seinen-shoujo 0.490
In other words, a shounen title resembles a shoujo title more than it resembles another shounen title. Same for seinen. Neither has its own label as the maximum of its row.
The reason is scale. Shounen's 260 titles are a third of the pool and include battle series, sports, gag comedies and isekai all at once. "Shounen" is not a kind of story — it is a note about the magazine's readership, so it does not map onto the plot structure similarity actually reads.
Shoujo's 42 titles, by contrast, are few and genuinely alike: relationships and feelings at the center, with repeating event structures.
3. How to use this in the game
- Do not narrow by "shounen feel." 0.49 is close to random. If a battle series landed, narrow to other battle series — not to other shounen.
- If relationship-driven signals appear, push into shoujo and josei. At 0.61 that side is a real cluster, so one hit pulls neighbors along.
- Anything near 0.5 means only "same demographic." You have the label and nothing else; switch axes to subject matter or tone.
- Themes beat demographics. The 34 theme tags spread from 0.643 to 0.853, while every demographic sits bunched between 0.47 and 0.61.
4. Limits
The matrix is not quite symmetric — josei-seinen reads 0.495 while seinen-josei reads 0.482. The similarity values themselves are symmetric, but each item stores only its top 3,000 neighbors, so some pairs appear in one list and not the other. Do not trust the third decimal; read only large gaps like 0.49 versus 0.61.
Josei (14) and kids (6) are small samples, and shoujo (42) is not large. The shounen (260) and seinen (108) conclusions rest on enough data, but the shoujo-josei cluster could shift with more.
Demographic labels do feed the structural text of the similarity calculation. That same-label pairs still only reach 0.49 means the difference in plot structure outweighs whatever the tag contributes. Method on the about page.
Further reading
A Sharper Axis Than Genre runs this measurement over 34 theme tags and Which Genres Are Hardest over 17 genres. Individual cases where tags disagree with similarity are in When Similarity Betrays the Tags. Each day's answer and its nearest neighbors accumulate in the past answers archive.