Measured results: XEYE accuracy and latency

Measured against the production service with 604 queries in Spanish: XEYE puts the expected element in first position in 91% of queries on a catalogue of 10 products, in 85% on 250 activities and in 49% on 4,727 near-synonymous codes, where it appears in the top ten 84% of the time. With the default model, the server answers in 28 to 55 ms on average.

By Joan Martorell

How was it measured?

Each test is a pair made of a query and the element it should return. Queries are sent to the public API with an API key, just as a real integration would send them, and we record the position at which the expected element appears. No query repeats the text of the element literally, because an exact match would measure nothing.

  • Twelve runs and 604 queries in total, in August 2026, against the production deployment.
  • Three catalogues of three sizes, each trained with the two available embedding models, with and without AI-generated descriptions.
  • Latency is the figure the service itself reports in each response: processing time on the server, without the network trip, with the list already loaded in memory.

The measurements are part of a bachelor's thesis at the Universitat de les Illes Balears (2026).

The metrics

  • Top-1: fraction of queries in which the expected element comes first.
  • Recall@k: fraction in which it comes in the top k.
  • MRR (mean reciprocal rank): the mean of 1 divided by the position of the expected element. It is 1 if the element always comes first.

Which catalogues were used?

CatalogueElementsQueriesWhat it contains
Products1035E-commerce products.
Activities25055Business activities, queried the way a user would describe their profession.
CPV4,72761Codes from the European Common Procurement Vocabulary: a nomenclature with dozens of near-synonymous codes per query.

Each query is labelled by what it tests: synonym (different words from those of the element), intent (it describes a purpose, not an object), brand, typo and attribute (colour, size, material).

What accuracy does it achieve?

CatalogueModelTop-1Recall@3Recall@5Recall@10MRR
Products (10)MiniLM0.911.001.001.000.95
Products (10)mpnet0.911.001.001.000.96
Activities (250)MiniLM0.850.960.980.980.91
Activities (250)mpnet0.850.981.001.000.92
CPV (4,727)MiniLM0.490.690.740.840.62
CPV (4,727)mpnet0.480.740.800.890.63
With AI descriptions turned on. MiniLM is the default model (384 dimensions); mpnet is the large model (768).

In the small and medium catalogues no query went unanswered: the expected element always appeared among the results, and quality barely drops when going from 10 to 250 elements. The CPV catalogue is a different regime: with thousands of near-synonymous codes, getting exactly the first one right is hard even for a person, but the expected element still appears in the top ten in 84 to 89% of queries.

How does it perform by query type?

Query typeProductsActivitiesCPV
Synonym0.89 (9)0.93 (29)0.43 (23)
Intent1.00 (8)0.77 (13)0.50 (26)
Brand0.83 (6)no queriesno queries
Typo1.00 (3)0.75 (4)0.67 (6)
Attribute0.89 (9)0.78 (9)0.50 (6)
Top-1 with the default model (MiniLM). The number of queries is in brackets.

Synonym and intent queries are the ones a keyword search engine cannot resolve. In the Products catalogue they hit the first position in 89% and 100% of cases, and in Activities synonym is the strongest category, while intent drops to 0.77. The reason why is explained in semantic search versus keyword search.

How much do AI-generated descriptions contribute?

To measure it, each list was trained again with AI descriptions turned off and the full evaluation was repeated.

CatalogueModelTop-1 with AITop-1 without AIMRR with AIMRR without AI
ProductsMiniLM0.910.860.950.92
Productsmpnet0.910.860.960.93
ActivitiesMiniLM0.850.690.910.80
Activitiesmpnet0.850.750.920.83
CPVMiniLM0.490.390.620.52
CPVmpnet0.480.430.630.54

Doing without them costs between 0.03 and 0.11 of MRR, and it costs more in the medium and large catalogues than in the small one. There is one exception worth knowing: in Activities, attribute queries improved without AI (from 0.78 to 1.00 top-1, over 9 queries). In very literal queries, the generated text can move the element slightly away from the words that describe it.

What is the latency?

CatalogueModelMean95th percentile
Products (10)MiniLM4470
Products (10)mpnet103174
Activities (250)MiniLM2832
Activities (250)mpnet8194
CPV (4,727)MiniLM5568
CPV (4,727)mpnet111133
Milliseconds of processing on the server, with the list already loaded.

Latency depends on the model far more than on the size of the list. It does not grow between 10 and 250 elements, and going up to 4,727 adds about 25 ms on average. Switching from MiniLM to mpnet, by contrast, multiplies it by 2 to 2.9 at any scale: what dominates is turning the query into a vector.

Which model should I choose?

In small and medium catalogues, the large model improves the metrics only marginally (0.01 of MRR) and does not make up for its latency. In the large catalogue it does help: with an almost identical top-1, it raises recall@3 from 0.69 to 0.74 and recall@10 from 0.84 to 0.89, and the mean position of the expected element goes from 5.1 to 3.4. The model is chosen when launching each training, so you can try both on your own list.

What are the limits of this data?

  • The test sets are small (35 to 61 queries) and were written by the author, not by real users.
  • All the queries and catalogues are in Spanish.
  • Some categories have very few queries: a difference between 0.75 and 1.00 over 4 queries is a single query.
  • Latency measures the warm case. The first search on a list after a service restart takes longer, and that case has not been measured systematically.
  • It is not a comparison with other products: none has been measured with these sets.

The best test is your own. The console playground is free: upload a sample of your catalogue and search the way your users would.

Frequently asked questions

What does a top-1 of 0.85 mean?

That in 85 out of every 100 test queries the expected element came in first position. In the rest it did not necessarily fail: recall@3 for that same catalogue is 0.96, so it was almost always in the top three.

Why does accuracy drop so much in the catalogue of 4,727 elements?

Because it is a hierarchical nomenclature with dozens of near-synonymous codes competing for each query. With a product catalogue of the same size but with elements that differ more from each other, the behaviour to expect is different, and we have not measured it.

Does the latency include the network?

No. It is the processing time the service reports in the duration_ms field of each response. To that you have to add the trip between your server and the API.

Try it with your own data

Create an account, upload a list and run your first search in about five minutes. You start with €5 of credit, no card required.

Keep reading