How does an AI software program remember a conversation you had two weeks ago while you are typing in an IDE today? Beneath the hood, AI memory relies on an elegant synthesis of vector embeddings, hybrid retrieval, and temporal graph theory.
1. Chunking and Preprocessing
When an AI memory assistant captures a meeting transcript or a document, it splits the text into semantic chunks (typically 256 to 512 tokens). Each chunk preserves speaker attribution, window titles, and exact timestamps.
2. Vector Embeddings: Giving Text Coordinates
Each chunk is passed through an embedding model (such as a quantized MiniLM or BGE model running locally). The model translates the words into an array of numbers — a vector — that captures its semantic meaning. Sentences with similar meanings produce vectors that are close together in high-dimensional space.
3. Hybrid Search: Vectors + BM25 Keywords
Vector search is brilliant at conceptual similarity, but it can struggle with exact keyword lookups (like finding a specific ticket ID such as JIRA-4921). Maple uses hybrid retrieval, combining:
- BM25 Keyword Search: Pinpoints exact identifiers, acronyms, and unique names.
- Cosine Vector Search: Finds conceptually related context even if the exact keywords differ.
- Reciprocal Rank Fusion (RRF): Merges the two result lists to surface the highest-confidence matches in under 50 milliseconds.
4. Temporal Context Filtering
Work changes quickly. A decision made on Monday might supersede a decision from last month. An advanced AI assistant that remembers your work weights search results by recency and project association, ensuring you receive up-to-date answers.
To learn more about local-first database performance, read our deep dive on SQLite as an AI document store and see how Maple's AI memory engine functions on your laptop.