Manticore Search supports both row-wise and columnar attribute storage . Row-wise storage works well when the working set fits in memory, but columnar storage becomes especially useful when large datasets exceed RAM: queries that need only a few attributes can read and cache mostly the data they actually use.
There was one important performance gap between the two layouts. During KNN vector search , Manticore rescored HNSW candidates using their original full-precision vectors. With the previous columnar access path, that step was much slower than with row-wise storage. In our DBpedia benchmark, row-wise storage delivered 2.80x–3.53x the KNN throughput.
The problem wasn't the columnar layout itself. Columnar vectors were read through a reusable buffer, which prevented Manticore from keeping stable pointers to multiple vectors and processing them in batches.
We changed columnar access to use memory mapping. This gives vectors stable addresses while still allowing the operating system to load and evict file pages on demand. The result is 2.57x–3.13x higher KNN throughput, reaching 85–92% of row-wise performance, without sacrificing columnar storage's ability to work efficiently when the data is larger than available memory.
How slow can rescoring from columnar storage be?
We measured KNN throughput on a 16-core AMD Ryzen 9 5950X using the DBpedia dataset:
- 975,000 vectors
- 1,536 coordinates per vector
- 1-bit quantization
- 5,000 different queries per run
- A working set that fits in available memory
- The default oversampling and rescoring behavior
The initial comparison used row-wise and columnar vector storage:
Across the three measurements, the row-wise path delivered 2.80x - 3.53x the throughput.
Graph traversal contributes equally in both results. The throughput gap comes from the rescoring pass.
Why the defaults make this important
Rescoring and oversampling are part of Manticore's default KNN behavior:
oversampling=3.0multiplies the requestedkbefore HNSW search, retrieving more approximate candidates than the final query needs.rescore=1fetches the original 32-bit vectors for those candidates, recalculates their distances, sorts them again, and returns the final topk.
The default query path is therefore:
requested k -> retrieve up to 3 x k candidates -> rescoring -> return k results
That changes how the benchmark values should be interpreted:
| Requested k | Target HNSW candidate pool | Final results after rescoring |
|---|---|---|
| 20 | 60 | 20 |
| 100 | 300 | 100 |
| 500 | 1,500 | 500 |
For example, k=500 makes HNSW search with an effective k of 1,500. Those candidates become eligible for exact rescoring, and the best 500 are returned. Filtering, disk chunks, and candidate availability can affect the exact number of physical reads, while the target candidate pool remains three times the requested result count.
Oversampling and rescoring improve ranking quality, particularly with quantized vectors. Disabling them changes the normal quality/performance tradeoff. Optimizing the default path benefits typical KNN queries directly.
The graph search is the same; the vector access is different
Row-wise and columnar KNN tables use the same HNSW graph search. In both scenarios, the graph navigates the approximate index and produces candidate document IDs. Vector storage becomes relevant after that stage, when rescoring needs the original full-precision value for every candidate.
The row-wise accessor can provide a stable address for each resident vector. Manticore can retain several vector pointers, prefetch their data, and calculate multiple distances together.
Columnar access works differently. It fetches a vector through a reusable read buffer. A later read can overwrite that buffer, so the rescoring code processes one vector before fetching the next instead of retaining a group of pointers for batch processing.
This distinction grows more expensive with high-dimensional vectors and as the user increases k. The benchmark illustrates the effect at effective candidate pools of 60, 300, and 1,500, but users can request other k values and the same mechanism applies.
Since HNSW traversal is the same, this rescoring behavior accounts for the observed storage-mode gap.
Why columnar storage still matters
Columnar storage was originally intended for cases where there is not enough memory to load all the data for a queried attribute. It arranges all values of one attribute next to one another. A page read for price, for example, contains mostly more price values rather than prices interleaved with categories, timestamps, and other attributes.
Row-wise storage keeps the attributes for one document together. That is useful when a query needs the complete row, while a query that reads one attribute may also bring unrelated values into memory. For scans, filters, and aggregations over individual attributes, this lowers the useful-data density of each loaded page.
The contiguous columnar layout normally behaves better under memory pressure because it reads and caches more of the requested attribute and less unrelated data. That advantage becomes especially important when the table is much larger than RAM.
The slowdown had a simpler cause: random vector reads during rescoring had to pass through one reusable buffer.
A new access path for columnar vectors
Memory-mapped file performance depends on the workload; for rescoring, the key benefit is stable access to the columnar vectors.
Mapping reserves virtual address space for a file. Physical pages enter RAM on demand as the process accesses them, and the operating system can reclaim those pages under memory pressure. A mapping can therefore be larger than the available physical memory.
For rescoring, pointer stability is the key change. A vector can be addressed directly in the mapped region instead of copied through a reusable read buffer. Manticore can then:
- Sort candidates by disk chunk and row ID to improve locality.
- Collect stable vector pointers in batches of up to 256 candidates.
- Prefetch vector data before it is needed.
- Calculate distances for the batch, processing pairs together where supported and handling any remainder individually.
Memory mapping enables the batched rescoring path that row-wise storage already benefited from.
Fit-in-memory result: most of the gap disappears
We repeated the DBpedia test with memory-mapped columnar access and compared all three storage paths:
At k=20, columnar throughput rose from 178 to 509 QPS, or 2.86 times the file result. At k=100, it rose from 97 to 304 QPS, a 3.13-times improvement. At k=500, it rose from 61 to 157 QPS, a 2.57-times improvement.
Columnar file access reached 28-36% of row-wise throughput. Memory-mapped columnar access reached 85-92%. Its remaining gap to row-wise storage was 15.4% at k=20, 11.1% at k=100, and 8.2% at k=500.
The narrowing gap is consistent with batching becoming more valuable as k and the resulting rescoring work increase. At the benchmark's k=500 point, memory-mapped columnar storage finished within about 8% of row-wise performance instead of delivering roughly one-third of its throughput.
That establishes the result for a resident working set. The next question is how the new path behaves in the memory-constrained conditions columnar storage was designed for.
Out-of-core result: the taxi benchmark
For generic search, we used a much larger taxi dataset on the same Ryzen 9 5950X:
- 1.74 billion documents
- 32 disk chunks
- 372 GB total table size
- 88 GB of queried
.spccolumnar storage files - 32 GB Docker memory limit
The queried columnar footprint was 2.75 times the container's entire memory limit. The search server also needed memory for its own data structures, leaving less than 32 GB for filesystem-backed pages. This is an out-of-core workload, so queries run under page eviction and physical I/O.
We ran the same generic-search suite against two versions of the data:
- taxi: the complete 1.74-billion-document table, which exceeds available memory.
- taxi1: one disk chunk, whose working set fits in memory.
The 17 queries included full-text search, unfiltered aggregates, equality and range filters, indexed lookups, and high- and low-cardinality GROUP BY operations. These are the same taxi queries used in our public comparisons on db-benchmarks.com
.
Each access path was tested in three complete runs. The figures below use the arithmetic mean of server-reported time across those runs, with minimum and maximum values shown by the whiskers. Each value represents the total for the entire 17-query suite.
Cold and hot measurements are analyzed separately:
- Cold is the first measured execution in the benchmark's cache-drop phase.
- Hot is the mean of 10 repeated executions per query on taxi and 50 on taxi1.
Cold runs
| Dataset | Columnar access | Average | Min-max | Change vs file |
|---|---|---|---|---|
| Whole taxi table, out of core | file | 25.086 s | 24.489-25.581 s | - |
| Whole taxi table, out of core | mmap | 25.155 s | 24.561-25.952 s | +0.28% |
| Single taxi chunk, resident | file | 845.667 ms | 793-905 ms | - |
| Single taxi chunk, resident | mmap | 832.667 ms | 815-848 ms | -1.54% |
Lower is better.
On the complete taxi table, mmap was 0.28% slower in cold runs, a negligible difference.
On the resident single chunk, mmap was 1.54% faster: 832.667 milliseconds versus 845.667 milliseconds, which is a small improvement.
Hot runs
| Dataset | Columnar access | Average | Min-max | Change vs file |
|---|---|---|---|---|
| Whole taxi table, out of core | file | 22.196 s | 22.033-22.521 s | - |
| Whole taxi table, out of core | mmap | 22.223 s | 22.074-22.500 s | +0.12% |
| Single taxi chunk, resident | file | 758.667 ms | 754.120-765.920 ms | - |
| Single taxi chunk, resident | mmap | 751.087 ms | 742.660-759.540 ms | -1.00% |
Lower is better.
On the out-of-core table, mmap was 0.12% slower in hot runs, a negligible difference.
On the resident chunk, mmap was 1.00% faster: 751.087 milliseconds versus 758.667 milliseconds. Grouping and range queries included improvements, while some simple aggregates moved slightly in the other direction. The mixed per-query changes and small aggregate difference again support near-parity.
Together, the cold and hot runs show that the mapped path preserves non-KNN search performance. It was neutral when the queried columnar files exceeded available memory and slightly faster when one chunk fit in memory.
Why non-KNN mmap performance remains close
The large KNN improvement comes from changing the rescoring strategy: stable vector pointers enable better locality, prefetching, and batched distance calculations. The non-KNN taxi queries do not use this batched KNN rescoring path, so they don't benefit from it.
On Linux, normal reads and file-backed mmap both go through the kernel page cache. The file mode reads cached data into an application buffer. mmap exposes the file-backed pages through the process address space and brings missing pages in through page faults. Both paths remain subject to the same physical-memory limit, page reclaim, and backing storage.
This shared foundation helps explain the non-KNN parity. mmap changes the pointer and access model enough to unlock batched KNN rescoring, while the operating system continues to manage the underlying file pages under both access modes.
Conclusion
Rescoring can be a substantial part of default KNN execution. Three-times oversampling means a query asking for 500 results uses an effective HNSW k of 1,500 before the rescoring pass. As users raise or lower k, the rescoring workload changes with it. When candidate vectors were read one at a time through columnar file access, KNN throughput was 64-72% lower than with row-wise storage, even though HNSW traversal itself was unchanged.
Stable mapped addresses allow Manticore to batch that work. On DBpedia, columnar throughput rose by 2.57-3.13 times and reached 85-92% of row-wise performance.
The memory-constrained tests show that the approach also preserves columnar storage's original strength. With 88 GB of queried columnar data under a 32 GB container limit, non-KNN search time changed by only +0.28% cold and +0.12% hot. The resident single-chunk results were slightly favorable at -1.54% cold and -1.00% hot. The overall result is a large improvement to default KNN rescoring with non-KNN search parity under both memory conditions. As of Manticore Search 29.9.0
, mmap is the default access mode for columnar storage.
Columnar access configuration
The access_columnar_attrs
option controls how columnar files are accessed. Its default value is mmap, so no per-table setting is required. The mode can also be selected explicitly when creating a table:
CREATE TABLE products (
title TEXT,
embedding FLOAT_VECTOR
KNN_TYPE='hnsw'
KNN_DIMS='1536'
) ENGINE='columnar'
access_columnar_attrs='mmap';
The previous file mode remains available. To use it as the server-wide default, add it to the searchd section of the configuration file:
searchd {
access_columnar_attrs = file
}
The option changes how columnar files are accessed. Their stored format and the KNN query syntax stay the same.

