Measured results: XEYE accuracy and latency
Measured against the production service with 604 queries in Spanish: XEYE puts the expected element in first position in 91% of queries on a catalogue of 10 products, in 85% on 250 activities and in 49% on 4,727 near-synonymous codes, where it appears in the top ten 84% of the time. With the default model, the server answers in 28 to 55 ms on average.
By Joan Martorell
How was it measured?
Each test is a pair made of a query and the element it should return. Queries are sent to the public API with an API key, just as a real integration would send them, and we record the position at which the expected element appears. No query repeats the text of the element literally, because an exact match would measure nothing.
- Twelve runs and 604 queries in total, in August 2026, against the production deployment.
- Three catalogues of three sizes, each trained with the two available embedding models, with and without AI-generated descriptions.
- Latency is the figure the service itself reports in each response: processing time on the server, without the network trip, with the list already loaded in memory.
The measurements are part of a bachelor's thesis at the Universitat de les Illes Balears (2026).
The metrics
- Top-1: fraction of queries in which the expected element comes first.
- Recall@k: fraction in which it comes in the top k.
- MRR (mean reciprocal rank): the mean of 1 divided by the position of the expected element. It is 1 if the element always comes first.
Which catalogues were used?
| Catalogue | Elements | Queries | What it contains |
|---|---|---|---|
| Products | 10 | 35 | E-commerce products. |
| Activities | 250 | 55 | Business activities, queried the way a user would describe their profession. |
| CPV | 4,727 | 61 | Codes from the European Common Procurement Vocabulary: a nomenclature with dozens of near-synonymous codes per query. |
Each query is labelled by what it tests: synonym (different words from those of the element), intent (it describes a purpose, not an object), brand, typo and attribute (colour, size, material).
What accuracy does it achieve?
| Catalogue | Model | Top-1 | Recall@3 | Recall@5 | Recall@10 | MRR |
|---|---|---|---|---|---|---|
| Products (10) | MiniLM | 0.91 | 1.00 | 1.00 | 1.00 | 0.95 |
| Products (10) | mpnet | 0.91 | 1.00 | 1.00 | 1.00 | 0.96 |
| Activities (250) | MiniLM | 0.85 | 0.96 | 0.98 | 0.98 | 0.91 |
| Activities (250) | mpnet | 0.85 | 0.98 | 1.00 | 1.00 | 0.92 |
| CPV (4,727) | MiniLM | 0.49 | 0.69 | 0.74 | 0.84 | 0.62 |
| CPV (4,727) | mpnet | 0.48 | 0.74 | 0.80 | 0.89 | 0.63 |
In the small and medium catalogues no query went unanswered: the expected element always appeared among the results, and quality barely drops when going from 10 to 250 elements. The CPV catalogue is a different regime: with thousands of near-synonymous codes, getting exactly the first one right is hard even for a person, but the expected element still appears in the top ten in 84 to 89% of queries.
How does it perform by query type?
| Query type | Products | Activities | CPV |
|---|---|---|---|
| Synonym | 0.89 (9) | 0.93 (29) | 0.43 (23) |
| Intent | 1.00 (8) | 0.77 (13) | 0.50 (26) |
| Brand | 0.83 (6) | no queries | no queries |
| Typo | 1.00 (3) | 0.75 (4) | 0.67 (6) |
| Attribute | 0.89 (9) | 0.78 (9) | 0.50 (6) |
Synonym and intent queries are the ones a keyword search engine cannot resolve. In the Products catalogue they hit the first position in 89% and 100% of cases, and in Activities synonym is the strongest category, while intent drops to 0.77. The reason why is explained in semantic search versus keyword search.
How much do AI-generated descriptions contribute?
To measure it, each list was trained again with AI descriptions turned off and the full evaluation was repeated.
| Catalogue | Model | Top-1 with AI | Top-1 without AI | MRR with AI | MRR without AI |
|---|---|---|---|---|---|
| Products | MiniLM | 0.91 | 0.86 | 0.95 | 0.92 |
| Products | mpnet | 0.91 | 0.86 | 0.96 | 0.93 |
| Activities | MiniLM | 0.85 | 0.69 | 0.91 | 0.80 |
| Activities | mpnet | 0.85 | 0.75 | 0.92 | 0.83 |
| CPV | MiniLM | 0.49 | 0.39 | 0.62 | 0.52 |
| CPV | mpnet | 0.48 | 0.43 | 0.63 | 0.54 |
Doing without them costs between 0.03 and 0.11 of MRR, and it costs more in the medium and large catalogues than in the small one. There is one exception worth knowing: in Activities, attribute queries improved without AI (from 0.78 to 1.00 top-1, over 9 queries). In very literal queries, the generated text can move the element slightly away from the words that describe it.
What is the latency?
| Catalogue | Model | Mean | 95th percentile |
|---|---|---|---|
| Products (10) | MiniLM | 44 | 70 |
| Products (10) | mpnet | 103 | 174 |
| Activities (250) | MiniLM | 28 | 32 |
| Activities (250) | mpnet | 81 | 94 |
| CPV (4,727) | MiniLM | 55 | 68 |
| CPV (4,727) | mpnet | 111 | 133 |
Latency depends on the model far more than on the size of the list. It does not grow between 10 and 250 elements, and going up to 4,727 adds about 25 ms on average. Switching from MiniLM to mpnet, by contrast, multiplies it by 2 to 2.9 at any scale: what dominates is turning the query into a vector.
Which model should I choose?
In small and medium catalogues, the large model improves the metrics only marginally (0.01 of MRR) and does not make up for its latency. In the large catalogue it does help: with an almost identical top-1, it raises recall@3 from 0.69 to 0.74 and recall@10 from 0.84 to 0.89, and the mean position of the expected element goes from 5.1 to 3.4. The model is chosen when launching each training, so you can try both on your own list.
What are the limits of this data?
- The test sets are small (35 to 61 queries) and were written by the author, not by real users.
- All the queries and catalogues are in Spanish.
- Some categories have very few queries: a difference between 0.75 and 1.00 over 4 queries is a single query.
- Latency measures the warm case. The first search on a list after a service restart takes longer, and that case has not been measured systematically.
- It is not a comparison with other products: none has been measured with these sets.
The best test is your own. The console playground is free: upload a sample of your catalogue and search the way your users would.
Frequently asked questions
What does a top-1 of 0.85 mean?
- That in 85 out of every 100 test queries the expected element came in first position. In the rest it did not necessarily fail: recall@3 for that same catalogue is 0.96, so it was almost always in the top three.
Why does accuracy drop so much in the catalogue of 4,727 elements?
- Because it is a hierarchical nomenclature with dozens of near-synonymous codes competing for each query. With a product catalogue of the same size but with elements that differ more from each other, the behaviour to expect is different, and we have not measured it.
Does the latency include the network?
- No. It is the processing time the service reports in the
duration_msfield of each response. To that you have to add the trip between your server and the API.
Try it with your own data
Create an account, upload a list and run your first search in about five minutes. You start with €5 of credit, no card required.
Keep reading
- Semantic search versus keyword searchKeyword search compares letters and semantic search compares meanings. When each one gets it right, what hybrid search is and measured data.
- XEYE pricingXEYE costs €0.001 per API search and €0.30 per training, plus €0.0053 per AI description. No fees, and €5 of starting credit.
- Semantic search API: what it is and how to choose oneA semantic search API returns the elements in your catalogue that mean the same as the query. What it does, what you need and how to choose one.