Ask Runable forDesign-Driven General AI AgentTry Runable For Free
Runable
Back to Blog
Technology8 min read

Announcing Evals and Releases: Evaluate Fin before, during, and after you go live - The Intercom Blog

Evals and Releases, together with Monitors, bring Eval-driven delivery to support teams, helping you maintain performance over time. Discover insights about ann

TechnologyInnovationBest PracticesGuideTutorial
Announcing Evals and Releases: Evaluate Fin before, during, and after you go live - The Intercom Blog
Listen to Article
0:00
0:00
0:00

Announcing Evals and Releases: Evaluate Fin before, during, and after you go live - The Intercom Blog

Overview

For customers Meet your customers where they already are with the world’s best business messenger for chat, email, voice, social…

Ideas blog Product & Design thoughts from our leadership team

Details

The Ticket podcast Conversations with future-focused leaders at the cutting edge of customer service.

For customers Meet your customers where they already are with the world’s best business messenger for chat, email, voice, social…

Ideas blog Product & Design thoughts from our leadership team

The Ticket podcast Conversations with future-focused leaders at the cutting edge of customer service.

Announcing Evals and Releases: Evaluate Fin before, during, and after you go live

Providing a complete evaluation system for Fin, you can now test changes before they go live, roll them out with control, evaluate every live conversation, and have confidence in the experience Fin delivers.

Maintaining customer experience as your business evolves

Today we’re announcing Evals and Releases. Paired with Monitors, they give you an evaluation system for Fin, so you can test changes before they go live, roll them out with control, evaluate every live conversation, and have confidence in the experience Fin delivers.

Once live, Monitors keeps the process continuous: assessing the quality of every conversation, and catching issues you can work through in your next Eval and Release.

We call this “Eval-driven delivery.” It’s the discipline our AI research team uses to build and tune AI products, and it’s now in the hands of the teams running support with Fin.

Evals tests Fin end to end. Create an Eval for a topic, add Simulations that mirror real customer conversations, and each one is scored pass/fail against your success and qualitative criteria. Re-run the Eval after any change to Fin to catch unexpected regressions in performance.

Evals, Releases, and Monitors work as one system. Monitors checks every live conversation against your standards and flags the ones that fall short. Turn any flagged conversation into your next set of improvements and run the system again.

As AI Agents take on more volume and complex work, ensuring they meet your standards every time is a tough task.

Fin, like any AI Agent, is probabilistic. It reasons through each conversation as it happens, so it can answer the same question in different ways – and your customers will ask the same question in countless ways too. That creates thousands of possible scenarios that no team can validate by hand.

And success isn’t just about whether the final answer was right. It depends on whether Fin followed your Procedures, used the right content and data sources, handed off at the right moment, and sounded like your brand while doing it.

Teams also continuously make changes to Fin’s content, updating its context when they launch a product, adjust a policy, rewrite a help article, or add a Procedure. That can introduce drift, where changing one thing can impact something else without you noticing. At the scale some of our customers run Fin, a 1% regression could affect thousands of conversations a day.

Current oversight and testing solutions weren’t built for this scale and complexity. Meaning, it can be really hard to always know how Fin will handle a conversation.

We think every support team should have the ability to rigorously test changes and monitor performance without needing a developer or ML specialist.

Evals: Know how Fin will behave before you go live

An Eval is a named group of Simulations – multi-turn test conversations – that validates Fin’s behavior on one theme, topic, or issue. You might build one for your most common queries, one for how Fin handles refund requests inside and outside your policy, and another for tone of voice across a range of situations.

The simulated customer – their opening message, and the context they reveal as the conversation goes on, like account details, order information, or a change in mood.

What Fin has access to – including attributes and data connectors.

The criteria the LLM judge scores against – whether Fin replied and what it needed to say, whether the right Procedure triggered, or whether a data connector was called.

You can create Simulations manually, upload them from existing conversations, or generate them based on real inbox conversations.

Score every Simulation the way your best reviewers would

Running an Eval runs every Simulation inside it and returns a pass or fail for each one. Scoring combines deterministic checks with an LLM judge, and if any single criterion fails, the Simulation fails.

Re-run an Eval after any change to catch regressions

Once an Eval exists, it becomes a way to test for regressions on an ongoing basis. Re-run it before or after committing a change to Fin so you can ensure a fix to one thing doesn’t break another.

“We group our Evals around different topics, so every time we make a change, we rerun the whole set and see immediately whether anything regressed.” — Hila Horenshtein, CX AI Operations Team Lead, Auto DS

“We group our Evals around different topics, so every time we make a change, we rerun the whole set and see immediately whether anything regressed.”

— Hila Horenshtein, CX AI Operations Team Lead, Auto DS

Releases: Control how a change reaches customers

Releases gives you a dedicated space to plan, collaborate, and test changes to Fin so nothing reaches customers until you decide it’s ready.

In a Release, you can edit content, add or update Procedures, and change or delete Guidance, all bundled together. This simplifies big changes like new product launches, or optimizations to a specific workflow – like moving Fin from explaining your refund policy to processing refunds with a Procedure.

You can also run an Eval against any changes within the Release before you publish it. Build the change, run your Eval, see what fails, make adjustments, and run it again – without touching live Fin.

When a Release is ready you can publish it to everyone or run an A/B test against Fin’s current configuration. Experiment reporting uses metrics you care about, like resolution rate, escalation rate and CSAT, so you can be confident in every new release.

If something looks wrong partway through a rollout, pause it. Fix the specific issue, re-run the Eval to confirm the fix worked, then pick the rollout back up. If you need to undo a Release entirely, roll it back in one step.

Each of these is useful alone. Together, they strengthen the Fin Flywheel, so only the best version of Fin reaches your customers.

You create a change in a Release and validate it with Evals before going live. You roll it out as an experiment, then publish to everyone once the results hold up. From there, a Monitor keeps watch – checking live conversations against your quality standards so the change keeps performing after it’s fully live. When a Monitor flags a conversation that fell short, you turn it into a Simulation, add it to the relevant Eval, and use it as a benchmark to test a new Release and drive better performance. That’s how you can have confidence in the experience Fin delivers at scale.

You don’t have to work through Evals and Releases step by step yourself. Operator, our Agent for customer operations, can use all of these capabilities on your behalf.

Update your refund policy to say that customer refunds can be applied as a credit on their account.

Then, add a test for this to the ‘refund policy’ Eval.

And it will handle all three. It drafts the content change, builds the Eval, and bundles the work into a Release. Every step comes back as a proposal, so you can review it and give final approval.

Maintaining customer experience as your business evolves

You can build a very good AI Agent on day one. Maintaining this performance while your products, policies, and content keep changing is a different problem.

Salesforce signs definitive agreement to acquire Fin

Announcing major updates to Procedures and Simulations: Enabling Fin to handle complex queries

Key Takeaways

  • For customers Meet your customers where they already are with the world’s best business messenger for chat, email, voice, social…
  • Ideas blog Product & Design thoughts from our leadership team
  • The Ticket podcast Conversations with future-focused leaders at the cutting edge of customer service
  • For customers Meet your customers where they already are with the world’s best business messenger for chat, email, voice, social…
  • Ideas blog Product & Design thoughts from our leadership team

Cut Costs with Runable

Cost savings are based on average monthly price per user for each app.

Which apps do you use?

Apps to replace

ChatGPTChatGPT
$20 / month
LovableLovable
$25 / month
Gamma AIGamma AI
$25 / month
HiggsFieldHiggsField
$49 / month
Leonardo AILeonardo AI
$12 / month
TOTAL$131 / month

Runable price = $9 / month

Saves $122 / month

Runable can save upto $1464 per year compared to the non-enterprise price of your apps.