Case study · Secret Sauce Labs
How Secret Sauce Labs runs continuous sentiment analysis for 3 cents per 100 sessions
3¢
per 100 sessions analysed for sentiment and intent
28-56x
more frequent runs - weekly became every 3 to 6 hours
2
workloads live: sentiment analysis and memory curation
“You don't always have to run everything realtime… I always think about what needs to be realtime, what can be run on Doubleword.”
- Deric Atienza, Full Stack Engineer, Secret Sauce Labs
Background
Users brief SecretSauce in threads, describing what they want to make and which marketing it is for. Those threads say more about user intent than any feature metric, so Secret Sauce Labs wanted to analyse them.
Problem
The challenge was extracting intent from hundreds of sessions a day, many with 10 to 50 messages per thread, while staying accurate and running the pipeline cost-effectively.
Secret Sauce Labs built the analysis on GPT models while prototyping. It worked well, but running it across every session at closed API pricing carried an ongoing cost that made it unsuitable for production. The team tested locally against 10 sessions at a time and went no further.
Solution
Secret Sauce Labs prototyped the feature on closed models, then moved it to an open model on Doubleword running on the flex (Async) tier. Nothing had to be migrated from a live setup.
That version ran on GitHub Actions, processing sessions once a week. It worked until the corpus grew, when processing thousands of messages ran past the two-hour GitHub Actions timeout.
The team moved the job to Multica, an agent orchestration platform, and switched to batch inference. This removed the timeout ceiling, and they ramped batches up to run every 3 to 6 hours.
For the integration, the team used the Doubleword provider plugin for Pi, so they never had to declare models and endpoints by hand. A custom adaptor keeps Pi and Multica working together consistently, setting the service tier for each run.
Secret Sauce Labs then extended Doubleword to memory curation, triggered by a webhook when a pull request merges. An agent reads the PR conversation, picks out the gotchas worth remembering - such as an agent debugging against dev when it should have used live data - and commits them to memory.
Results
-
3 cents per 100 sessions - continuous sentiment and intent analysis across the full corpus.
-
28 to 56 times more frequent - a weekly run became one every 3 to 6 hours, once batch inference removed the runtime ceiling.
-
The full corpus, not a sample - every session is analysed on every run, rather than the subset a prototype could cover.
-
Two workloads made viable for production - sentiment analysis and memory curation both run on Doubleword.
Doubleword removes the default of using realtime and treating all requests as if they are the same, and offers comparable intelligence to closed models at a fraction of the price.
“If we didn't know about Doubleword, we would have thought we'd had to run it through an API key, which would just eat up our credits and tokens.”
- Deric Atienza, Full Stack Engineer, Secret Sauce Labs
What's next
Next, Secret Sauce Labs is building evaluations to ensure consistently good service for customers and to catch issues in agentic flows before they hit production, with Doubleword already chosen as the inference layer to drive them.
Beyond that, the team is using a blend of open and closed inference to develop new features and handle operations. Many tasks can be routed to open models, freeing budget for when a frontier closed model actually makes sense.
