Here's the dirty secret of AI scoring: the number always looks the same. A "74% fit" renders just as confidently whether it was computed from a deep project portfolio and a successful AI analysis, or from three projects and a forty-word scraped posting. Most tools never tell you which one you're looking at. We think that's the single biggest reason experienced principals don't trust AI scores — and they're right not to.
Every opportunity in your ArchFP discovery feed carries a Fit score from 0 to 100, personal to your firm. It's built from your actual portfolio — the sectors you really build in, learned from your Firm Projects rather than a settings form — and an AI pass separates the look-alikes a keyword match can't: for a school-focused practice, a school beats a transit depot, even when the depot is the bigger job. Sort by Fit and the work you're positioned to win rises to the top, each row carrying a one-line reason.
And if your firm genuinely does two kinds of work — schools and labs, say — the score now reads each opportunity against the right side of your practice, instead of a blurred average of both.
These past few weeks we shipped the layer we think actually matters. Every Fit score now carries a confidence level, computed from the evidence behind it: how many past projects the comparison rests on, how much substance the RFP posting actually had, and whether the deeper AI analysis ran or fell back to basic matching.
When the evidence is thin, the score says so. It renders muted, tagged "est.", and the tooltip tells you exactly why: estimate — based on 2 past projects · short RFP description. A strong score on strong evidence looks like a verdict. A score on thin evidence looks like what it is: an estimate.
A confident-looking number built on thin evidence isn't a small flaw in an AI tool. It's the whole reason experienced people stop trusting them.
Your team's behaviour is the best ground truth there is — what you save, what you pursue, what you win, and what you dismiss. Once your firm has a real track record on the platform, those signals begin to refine the ranking: opportunities similar to the ones you pursue get a nudge up; opportunities similar to the ones you dismissed get a nudge down.
We built this carefully. The behaviour signal doesn't switch on until your team has enough recorded actions to mean something. Dismissing an obvious mismatch teaches the system nothing — but dismissing something that scored well is treated as the valuable correction it is. And the signal refines the score; it never dominates it.
No portfolio yet? No fake number. A brand-new account sees "Add past projects to unlock Fit scoring" — because a fit score with nothing to compare against would be theatre, not signal.
No silent degradation. If the deeper AI pass can't run, you still get the deterministic score — and the confidence drops, with the reason shown. The system never quietly gets worse without telling you.
We hold the score accountable. We track whether high-fit opportunities are actually the ones firms end up pursuing, so the scoring is calibrated against reality — not vibes.
Back in July we wrote about teaching our review AI to say "I'm not sure". This is the same conviction applied to discovery: in an industry where proposals carry real money and real reputations, an AI score you can't interrogate is a liability wearing a badge. One that shows its evidence, admits its limits, and learns from your judgment is something closer to a colleague. That's the bar we're building every ArchFP feature to.
If you'd like to see your firm's own feed ranked — complete with scores that admit what they don't know — book a 15-minute demo and we'll run it on your firm's real portfolio.