# Specific Labs

> Source: https://aiwiki.ai/wiki/specific_labs
> Updated: 2026-09-15
> Fact-checked: 2026-09-15
> Categories: AI Benchmarks, AI Companies, Data & Datasets, Enterprise AI
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Specific Labs." aiwiki.ai, 15 Sept 2026. https://aiwiki.ai/wiki/specific_labs
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

Specific Labs is a San Francisco company that buys and licenses operational data and source code from businesses, and packages it as training and evaluation material for AI labs. Its Y Combinator entry describes the product as "Real World Environments and Data for AI" and says the next frontier of AGI is bottlenecked by data that is hard to access; the company site puts the same point as needing the decisions, exceptions, handoffs and judgment behind how work actually gets done inside a company.[1][3] It went through Y Combinator's Fall 2025 batch and is best known for Real-SWE, a September 2026 benchmark that scored frontier coding agents on private production codebases.[3][4]

## Founding and team

The company was founded in 2025 by Janak Sunil, its chief executive, and Siddhant Paliwal, its chief technology officer. Sunil was previously at Coinbase. Paliwal was previously at Third Chair, a Y Combinator X25 company, and at Intel. As of September 2026 Y Combinator listed the company as active, based in San Francisco, with a team of six and Pete Koomen as its primary YC partner.[3] The Real-SWE report carries a third byline, Snagnik Das, alongside the two founders.[4]

Y Combinator's directory records an earlier launch by the same company, Bear, described as a tool for tracking and influencing how often a brand appears in [ChatGPT](https://aiwiki.ai/wiki/chatgpt), [Claude Code](https://aiwiki.ai/wiki/claude_code), Perplexity and similar assistants.[3] The company's current contact link is a Calendly page under the handle janak-usebear.[1]

## What the company sells

Specific's pitch is directed at two sides of a market. To frontier labs and enterprises, it offers datasets and environments "representative of work inside real-world businesses."[3] To companies that hold such records, it offers money: a partnership page tells prospective sellers their company data "could be worth $100K-$1M" and collects details of headcount, years in operation and which tools the data lives in for a private fit review.[5] A separate referral program pays up to $10,000 to whoever is first to introduce a company with which Specific goes on to complete a data purchase.[6]

Its distinguishing claim is provenance: not synthetic tasks or open-source repositories, but licensed records of work that employees were paid to do. The Real-SWE launch post asked companies willing to contribute a codebase to get in touch, saying that Specific handles anonymization and that the resulting tasks stay private.[4]

## Benchmarks

Specific publishes benchmark reports under the same brand, testing models against what it calls human-verified data drawn from real conversations and real work rather than curated test sets.[7] Two are public as of September 2026, with a third listed as in progress.

| Report | Published | Subject |
| --- | --- | --- |
| bench-1, Speech-to-Text on Real Multilingual Conversation | June 2026 | 15 transcription models across Spanish, Gulf-dialect Arabic, Hinglish and Japanese, graded against human-verified transcripts |
| bench-2, Real-SWE | September 2026 | 8 model-and-harness pairs on 10 tasks from private production codebases |
| bench-3 | in progress | not announced |

The speech report found ElevenLabs Scribe v2 strongest overall and the only model handling code-switching well, and reported that no model transcribed Gulf-dialect Arabic better than roughly a 53% word error rate, meaning the best of them missed about every other word.[8]

## Real-SWE

[Real-SWE](https://aiwiki.ai/wiki/real_swe) is the company's software-engineering benchmark, released in September 2026. Ten tasks drawn from licensed private codebases were given to eight model-and-harness pairs over 640 rollouts. [Claude Fable 5.1](https://aiwiki.ai/wiki/claude_fable_5_1) in Claude Code led at a 38.8% resolution rate and GPT-5.6 Sol in Codex CLI came last at 16.2%, with six of the ten tasks resolved less than 15% of the time and one resolved by no model at all.[10]

Specific announced it as a Y Combinator launch on September 10, 2026. Two days later it was submitted to Hacker News, where it had drawn 274 points and 156 comments by September 15, 2026, a mixture of practitioners reporting that the ranking matched their experience and readers objecting that a benchmark nobody outside the vendor can rerun is not verifiable.[9] Sunil answered questions in the thread, stating that the licensed codebases were written before 2023, that the company vets each contributing company by hand, that all models ran at high reasoning effort inside their provider's native harness, and that Specific intends to open-source some tasks and model trajectories.[9]

The launch post also invited labs that wanted "the full task set for training" to email Sunil directly, which makes the report both a published evaluation and a sales document for the underlying dataset.[4]

## What is not public

Little about the company's finances or customers is disclosed. Its site and its Y Combinator profile name Y Combinator as its only listed backer, give no funding amount, valuation or investor list, and name no lab or enterprise customer.[1][3] The companies whose codebases appear in Real-SWE are described only by category and scale, such as an events app with more than 200,000 users and a consumer fintech platform processing more than 100,000 bank statements, and Specific makes the full task set available only on request, by email to labs that want it for training.[4][10]

## References

1. [Specific Labs](https://withspecific.com) - company site.
2. [About](https://withspecific.com/about) - Specific Labs.
3. [Specific Labs](https://www.ycombinator.com/companies/specific-labs) - Y Combinator company directory.
4. [Real-SWE: A coding benchmark built from private company codebases](https://www.ycombinator.com/launches/TpS-real-swe-a-coding-benchmark-built-from-private-company-codebases) - Specific Labs launch post, Y Combinator, September 10, 2026.
5. [Company data partnerships](https://withspecific.com/company-data) - Specific Labs.
6. [Refer a company](https://withspecific.com/refer) - Specific Labs.
7. [Benchmarks](https://withspecific.com/benchmarks) - Specific Labs.
8. [Speech-to-Text on Real Multilingual Conversation](https://withspecific.com/benchmarks/speech-to-text) - Specific Labs, June 2026.
9. [Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases](https://news.ycombinator.com/item?id=49676820) - Hacker News discussion, submitted September 12, 2026.
10. [Introducing Real-SWE](https://withspecific.com/benchmarks/real-swe) - Snagnik Das, Siddhant Paliwal and Janak Sunil, Specific Labs, September 2026.
