Case study 6

Lead finding on Reddit: 96 real leads where Opus 5 found 72

ibo swapped JEV in as the lead finder inside peeklens.ai and ran it against gpt 5.6 sol and claude opus 5 over the same 1,511 Reddit posts with the same prompt. JEV returned more real leads (96 vs 78 vs 72), kept far less junk (4 vs 38 and 50), and cost $0.03 per scan against $0.20 and $0.28, under a second per post.

Original post "jev is now the lead finder in peeklens.ai ran it against gpt 5.6 sol and claude opus 5, same 1,511 reddit posts, same prompt - more real leads: 96 vs 78 vs 72 - 4 junk kept vs 38 and 50 - $0.03 per scan vs $0.20 and $0.28 - under 1s per post first real jev use case for indie devs" Read the original post on X

The case

  • peeklens.ai is a tool that finds leads for indie developers.
  • Lead finding was rebuilt so JEV is the model making the call: is this Reddit post a real lead or not.
  • All three models got the identical input: same 1,511 Reddit posts, same prompt. No prompt tuning in JEV's favour, per the post.
  • Two things were measured: how many real leads came back, and how much junk came with them.
  • Junk kept is the number that moves a workflow. 4 false positives against 38 and 50 is the difference between a list you work through and a list you clean first.

The numbers

metricJEVgpt 5.6 solclaude opus 5
real leads found967872
junk kept43850
cost per scan$0.03$0.20$0.28
speedunder 1s per postnot reportednot reported
metricreported
posts scanned1,511
promptidentical across all three
engagementnot captured in this pass

A reply from @SlimAssiliX frames why the numbers matter together: "96 real leads versus 72 at $0.03 versus $0.28 per scan means jev's precision advantage compounds with volume, the cost gap widens faster than the quality gap closes."

Why JEV is the right model here

Lead finding is a filter, not a conversation. Every post needs the same judgment: is this a buyer, or is this noise. That is one question asked 1,511 times, which is exactly the shape a decision model is built for.

A generative model answers that question by writing out its reasoning for each post, and you pay for every word of that reasoning. Here the reasoning was not the deliverable. The verdict was. JEV returns the verdict and a confidence score, which is why the scan costs cents and finishes in under a second per post.

The precision gap also compounds in a way the cost gap does not. Junk kept is manual labour handed back to you: 38 or 50 false positives means someone reads and discards them, and that work scales with every scan you run.

One honest caveat

The comparison is the author's own run, not an independent one, and "real leads" and "junk" are his labels for the outputs, so the bar was his to set. The prompt is described as identical across all three models, which makes this a clean prompt-for-prompt comparison, but prompt parity against models executing the same instruction is still not the same as a tuned setup per model.

The engagement numbers on the post itself were not captured this pass, so treat the result as reported-at-source rather than verified by us.

How to copy it

  1. Pick the repetitive judgment you already make by hand, in this case: is this Reddit post a lead.
  2. Define the outcomes as fixed options, not open ones. Lead or not a lead, with a confidence score on each.
  3. Run your existing model and JEV over the same batch with the same prompt, so the only variable is the model.
  4. Measure the two things that change your workload: how many real leads you got, and how much junk you have to clean.
  5. Compare cost per scan, then decide. The cost is what makes it viable to run continuously instead of occasionally.

Want this set up for your business?

We turn these workflows into working marketing and sales systems.

Book a call