Yuki Tanaka’s model-by-model source counter Prepared by Yuki Tanaka

Model Source Room

Where citation preferences become a commercial variable.

A grounded guide to how LLMs choose, ignore, reuse, and remix sources across retrieval systems and training influence, so teams can tune evidence for the model they actually need to persuade.

Counter premise

Models do not cite “the best source” in the abstract. They select from habits.

Model Source Room reads retrieval, citation, and training-data influence as source preference: what a model is willing to quote, what it paraphrases without naming, what it distrusts, and what it surfaces only when a prompt forces the issue.

Four readings per source

A source is judged by condition, not just relevance.

01

Retrieval pull

Whether the source is likely to enter the candidate set before the model begins answering.

02

Citation lift

Whether the source is named, linked, summarized, or quietly absorbed into the response.

03

Training echo

Whether older learned patterns overpower the retrieved document sitting in front of the model.

04

Commercial fit

Whether the source helps a buying team win a precise answer, not merely a mention.

Featured source pour

This week: the same document is not the same source to every model.

We compare source appetites — vendor docs, benchmark pages, analyst pages, community threads, PDFs, dated articles, and structured references — then translate those patterns into publishing and RAG decisions.

Source-condition bins

The room separates source visibility from source usefulness.

Quoted quickly
Easy to retrieve and easy for the model to defend.

Read but unnamed
Useful context that fails to become a visible citation.

Overridden by memory
Retrieved evidence that loses to training-era assumptions.

Useful only in a stack
Sources that need schema, corroboration, or adjacent proof before they lift.

Archive

Recent source flights