How a $180 claim becomes a $54 expectation.
We ask you not to take a vendor’s number on faith. It would be hypocritical to then ask you to take ours. This is the framework behind every adjustment we make — written out, so you can disagree with it specifically rather than generally.
Six questions asked of every claim.
None of these assume bad faith. A vendor reporting a real result from a real study is behaving reasonably — the problem is that the result was produced under conditions that aren’t yours, and nobody adjusts for the difference before the number reaches your desk.
One test comes before all six, costs ninety seconds, and settles a surprising number of claims on its own: check whether the vendor’s own number is arithmetically possible before opening the study behind it.
Every one of them is written out, applied and totalled in a full worked report — including the section that lists every assumption behind every number, which is the only part of this that proves the rest.
Four of the six are dials — you can move them yourself on the interactive report and watch the answer respond. The other two come out of reading the vendor’s actual study, which is the honest difference between the free tool and an engagement.
Selection bias
Who was actually in the study?
Programs are usually evaluated on the people who enrolled — and people who enroll in a musculoskeletal program are people who already decided to do something about their back. They were going to improve at a higher rate than the general population regardless. Unless the study used a matched comparison group drawn the same way, some of the reported effect belongs to the enrollment decision, not the program.
Typically removes 30–40% of a claimed effect.
Double-counted value
Who else is already being paid for this outcome?
A new program rarely lands on empty ground. Care management, the PBM's clinical programs, the health plan's own outreach and an EAP may all touch the same member and the same claim. Each vendor counts the full saving, correctly, in isolation. Added together, the value claimed across a portfolio routinely exceeds what was ever there to claim — and no single program is positioned to see that happen.
Typically removes another 15–25%.
Evidence transfer
Does this result move to your workforce?
A result generated in a 40-year-old professional services population doesn't transfer cleanly to a 52-year-old manufacturing workforce with different injury patterns, different scheduling and different care-seeking behaviour. We score the distance between the study population and yours across several dimensions and discount accordingly.
Varies most — the single biggest driver of the adjusted figure.
Engagement realism
What happens if fewer people show up?
Almost every savings figure is quoted at an engagement rate the vendor achieved somewhere. It's the assumption vendors are least specific about and the one that moves the answer most, so we never accept it as fixed — we model across a range and show you where their number sits inside it.
Shown as a range, never a point estimate.
Secular trend
Was this already happening without anyone?
A category can improve on its own. Surgical rates fall, a drug goes generic, a national care pattern shifts. A program running through that period will show a saving whether or not it caused one, so we take out the movement the category was making anyway. Where a category's trend was already negative nationally, no program gets credit for matching it.
Removes whatever the market was doing without you.
Verifiability
Was there anything to compare against?
Some claims rest on a before-and-after inside the enrolled group with no comparison group at all. That is rarely dishonest — it is often all the vendor had — but it means the figure carries no evidence about what would have happened otherwise. Where nothing can be compared, we report the number as unverifiable rather than discounting it to a smaller number that would look more precise than it is.
Reported as unverifiable, never estimated.
Every one of these adjustments is visible in your report, with the figure we used and where it came from. Disagree with one and the model re-runs.
What actually runs when we look at a vendor’s claim.
Six research passes run against your intake — four at once, then two more — before anything reaches the attribution model. The pipeline ends on human review, which is the point: nothing is published automatically.
- INTAKEYour programs, workforce and the vendor's claim
- WAVE 1 — 4 IN PARALLELCompany profile · Benefits scan · Financial signals · Workforce data
- WAVE 2Regulatory check · Workforce economics
- ATTRIBUTION MODELSelection bias, double-counted value, evidence transfer
- HUMAN REVIEWRead and edited by a person before release — every time
- YOUR REPORTRanges, not point figures, with every assumption shown
Everyone advising you gets paid. Almost none of them get paid less if you overspend.
A method is only worth as much as the independence behind it. You have no shortage of people telling you things about your benefit programs — it’s worth knowing what happens to each of their revenue when your spend goes up.
Closing the sale. Their outcomes study is, quite reasonably, sales collateral.
Serving you well — inside a structure that happens to pay more when you spend more.
Showing you the data. Deciding what to do about it is still your job.
Being right. That's the only thing we're paid for.
None of this says anyone is behaving badly. Plenty of brokers are excellent, fee-based advisors don’t carry this conflict at all, and a vendor publishing a real result from a real study is doing what any company would. We work alongside brokers and consultants regularly — several commission us directly.
The point is narrower than that: nobody in this picture is paid specifically to check the number.
A gym membership and an MSK program compete for the same dollar.
They are never compared, because they aren’t sold in the same units. Point solutions are quoted in avoided claims. Perks are quoted in retention and attraction — or not quoted at all. Two currencies, one budget line, and no exchange rate.
Nobody in the advice chain crosses that line, and not because anyone is failing at their job. Perks are largely unbrokered: there’s no commission, no catalogue entry and no reason for the question to reach the table. A musculoskeletal vendor cannot recommend a fitness stipend instead of itself.
Every program in our library — clinical and not — carries the same four scores: perceived value, financial leverage, retention, and clinical impact. That is what allows a stipend and a point solution to appear in one ranking rather than two conversations.
High-paid, hard-to-replace workforces have low claims utilisation relative to compensation. The clinical case is weakest precisely where the retention case is strongest — so this is the population where the standard stack is most likely to be the wrong answer.
Under a retention-weighted objective set we will tell you the stipend outranks the fourth overlapping point solution. We will not tell you it is worth forty dollars a month in retention, because nobody honestly can. See below.
What this method can’t tell you.
A framework that claims to answer everything is the same kind of overreach we exist to catch. Here’s where ours stops.
We can tell you what a program is likely worth under stated assumptions. We cannot tell you what it will return. Anyone who says otherwise is selling something.
Satisfaction, retention and productivity effects are real but too confounded by compensation, management and the labour market to attribute to a benefit decision with any honesty. We’ll discuss them; we won’t put a number on them.
Our analysis informs a decision; it doesn’t certify a reserve or clear a compliance question. Where something needs an actuarial opinion or legal review, we’ll say so rather than approximate it.
The questions we actually get asked.
Isn't this what my broker already does?
Partly — and if your broker is good, we'll say so. The difference is compensation. Most broker and consultant revenue rises as your benefit spend rises, which makes an honest recommendation to spend less structurally difficult. We take no commission from any vendor, carrier or broker, so the only thing we're optimising is whether the number holds up. We work alongside brokers regularly; we're not a replacement for one.
What benchmark data are you comparing me to?
Today: published industry survey data, public filings and regulatory sources, vendor-published outcomes, and two decades of prior modelling work across payer, provider and pharma. That's enough to place a portfolio credibly, and we'll tell you plainly when a comparison is directional rather than precise. Over time the benchmark becomes proprietary — every relationship adds to it — but we'd rather understate what we have now than overstate it.
Do I have to switch anything or install software?
No. There's nothing to roll out, no data feed to build, and no integration. You send us what you already have — vendor decks, renewal packets, benefit summaries — and we send back an analysis. If you never talk to us again, nothing breaks.
What if the answer is that our programs are fine?
Then that's the report. A finding of 'nothing to change here' is worth as much as a finding of savings — it's what stops you renegotiating something that's already working, or replacing a vendor who's performing. We'd rather be useful than dramatic.
How can you analyse our programs without our claims data?
Most of what determines whether a vendor's claim holds up isn't in your claims file — it's in the study design, the contract terms, the overlap with what you already run, and your workforce composition. Those come from documents you already have. Claims-level analysis is a deeper level of the relationship and needs a secure path we'll set up separately.
Why is the report free? What's the catch?
It's how we meet people, and it's the fastest way for you to judge the work without a sales process. There's no call attached to it and no obligation afterwards. If it's useful, the ongoing relationship is there; if it isn't, you've lost twenty minutes and gained a benchmark.
Run it on something you’re actually being sold.
The fastest way to judge a method is to point it at a claim you already have on your desk. Free, reviewed by a person, back within 24 hours.