Straight answer
Pick one or two real data sources, agree success criteria from the criteria that matter to you, then test four things: time to the first hunt, one advisory from report to tested rule, what a continuous hunt hands an analyst, and what data leaves your platforms. Score the results on the same scale for every tool.
Why run one at all
Every score on this site comes from public vendor material. A proof of concept is where a buyer checks those claims on their own data. It is also where the rows our reasons mark "not described" or "not stated" get an answer.
Step 1: pick the scope
Choose one or two real data sources that matter to your hunts, for example your EDR telemetry and your identity provider logs, or your SIEM and one cloud account. Fewer sources, tested properly, tell you more than a long list connected in a hurry. Agree a fixed window with each vendor before you start, and use the same window for every tool.
Step 2: write the success criteria first
Take the requirements you wrote from our criteria (see Writing requirements from the seven criteria) and turn each into something you can observe during the trial. Agree in writing what counts as a pass. Share the criteria with each vendor so nobody is surprised.
Step 3: time the path to the first hunt
Start the clock when access is granted. Stop it when the first hunt returns results from your own data. Write down every setup task on the way, who did it and whether any of it was a new pipeline, a forwarding rule or a new copy of data. This is the practical test of the time to first value row, which carries 15% here.
Step 4: run one advisory from report to tested rule
Pick a recent public advisory relevant to your estate. Give it to each tool and watch the path: reading the advisory, extracting behaviours, writing a query in each platform's own language, testing it against your past data, and deploying it. Record which steps were automatic, how the test was run and on how much data. Mars Security, for example, states a 30-day backtest on the customer's own data; Nebulock requires every rule to pass a retrohunt before deployment. Check what you see against what is stated.
Step 5: let a hunt run without you
Leave the tool to run continuous hunts for part of the window. At the end, look at what reached the analyst: a list of events, a written finding, a case or a proposed rule. Check whether each step is shown so an analyst can follow the reasoning, and how many findings needed no follow-up.
Step 6: check what left your platforms
Ask each vendor for a written list of any data copied out of your platforms during the trial, where it is stored and for how long. Compare it with your own logs where you can. Ask your platform owners whether query load changed. This is the practical test of the data reach row, which carries 25% here.
Step 7: check the rule lifecycle
Pick one rule the tool created and look at its history: who changed it, when, what the test showed, and how to roll it back. Try exporting it to your own repository if the vendor says that is supported.
Where do proofs of concept go wrong?
Three patterns make results hard to compare. Tools are tested on different data, so one looks better only because its sources were cleaner. Success criteria are set after the trial, so they drift toward whatever the favourite did well. And vendor staff run the trial on the buyer's behalf, so setup effort is hidden. Use the same sources, write the criteria first and have your own team do the setup steps wherever the vendor allows it.
Step 8: score and compare
Score each tool on the same 1 to 5 scale for each criterion, with the evidence next to each score. Apply your weights in the calculator to compare them with our reading of public material. Where your result and ours differ, your result wins: it was measured on your data.
Sources
- https://securityboulevard.com/2026/09/mars-security-launches-real-time-intel-to-detection-engine-that-turns-live-threat-intelligence-into-backtested-detections-in-minutes/
- https://docs.nebulock.io/docs/detection-engineering
- https://threathuntingcompare.com/how-we-compare
Editorial assessment · Desk research from public vendor material, last reviewed September 2026