Back to InspirationData

    RAG on company data: when to embed and when to just query

    September 8, 20269 min read
    RAG on company data: when to embed and when to just query

    The most common mistake in AI data projects is embedding everything. Vector search is for unstructured language. Company records are structured, and structured data should be queried, not retrieved by similarity.

    Draw the line

    Fields like registration number, status, sector code, employee band, incorporation date and address belong in a relational store with filters and aggregates. Similarity search on a size band will happily return the wrong band.

    Descriptions, news articles, filings text and reviews are language. Those are the right candidates for embeddings.

    A hybrid pattern

    Filter first in SQL — country, sector, size, status — then run vector search inside that filtered set, then let the model summarise the top results with citations. Answers stay correct because counting and filtering never touch the model.

    Chunk by document, keep the company identifier in the metadata of every chunk, and always return that identifier so any claim can be traced back.

    Evaluate on real questions

    Write thirty questions your team actually asks and score answers on correctness and citation, not on tone. Rerun that set on every prompt or model change.

    Watch for numeric questions: 'how many companies in this region' should always be answered by a query, never by a language model reading chunks.

    Frequently asked questions

    Do I need embeddings at all for company data?

    Only for the language parts: descriptions, news, filing narratives. Identity and firmographics are better served by filters and exact lookups.

    How fresh does the index need to be?

    Match it to the volatility of the field. Status and address changes matter within days; sector codes and incorporation dates almost never change.