jargon

Comparison

ChunkingvsSemantic chunking

Chunking

you split the document into pieces before you store it, because embedding a whole 40-page PDF as one blob retrieves nothing useful.

Splitting documents into pieces before embedding, because embeddings represent one 'thought' well and a 40-page document poorly, and because retrieved chunks must fit a context budget. Chunking quality caps retrieval quality; garbage chunks, garbage answers.

Full entry →

Semantic chunking

you split on headings and paragraph boundaries instead of every five hundred tokens, and retrieval on your documentation gets better.

Splitting on meaning boundaries (headings, paragraphs, topic shifts) instead of fixed token counts. More work, usually better retrieval, especially for structured documents like documentation and contracts.

Full entry →

Related comparisons