Enterprise

Deepset raises $14M to help companies build NLP apps

Comment

Image Credits: Getty Images

Natural language processing (NLP), the field of AI that involves parsing text for tasks including summarization and generation, is a fast-growing technology. According to a 2021 survey from John Snow Labs and Gradient Flow, 60% of tech leaders indicated that their NLP budgets grew by at least 10% compared to 2020, while a third said that their spending climbed by more than 30%. Fortune Business Insights pegged the NLP market at $16.53 billion in 2020.

Against this backdrop, Deepset, the startup behind the open source NLP framework Haystack, today announced that it raised $14 million in a Series A investment led by GV with participation from Harpoon Ventures, System.One, Lunar Ventures and Acequia Capital. The capital infusion arrived alongside Deepset Cloud, a new subscription product for building NLP-powered software.

“Driven by [our] belief in open source, the Deepset team has … been contributing models and research outcomes to the open source NLP community [for years],” Rusic told TechCrunch via email. “Haystack, the company’s flagship open source product, was born out of the experiences, expertise and know-how gained while building NLP for large organizations and the need for a proper set of building blocks for scalable, API-driven NLP back-end applications.”

CEO Milos Rusic co-founded Deepset with Malte Pietsch and Timo Möller in 2018. Pietsch and Möller — who have data science backgrounds — came from Plista, an adtech startup, where they worked on products including an AI-powered ad creation tool.

Haystack lets developers build pipelines for NLP use cases. Originally created for search applications, the framework can power engines that answer specific questions (e.g., “Why are startups moving to Berlin?”) or sift through documents.

Haystack can also field “knowledge-based” searches that look for granular information on websites with a lot of data or internal wikis. Rusic says that Haystack has been used to automate risk management workflows at financial services companies, returning results for queries like “What is the business outlook?” and “How did revenues evolve in the past years?” Other organizations, like Alcatel-Lucent Enterprise, have leveraged Haystack to launch virtual assistants that recommend documents to field technicians.

Haystack
A screenshot of the Haystack interface. Image Credits: Haystack

According to Rusic, the goal with Haystack was to enable developers and product divisions to build modern, API-driven NLP apps successfully — and quickly. He notes that, while it’s often straightforward for a data science team to come up with a prototype, challenges can arise in transitioning from prototype to production. About 80% of AI projects — including NLP projects — never make it into production, according to a 2019 Gartner survey.

“[With Haystack,] development teams … are equipped with all the components to build a full-stack NLP application and are guided with the proper workflows … Modern NLP moves very fast, and it’s much easier to bridge the gap between the cutting-edge research and the actual production-ready technologies through open source,” Rusic said. “[Prebuilt NLP systems] are the basis [for Haystack] and often provide great results in pipelines without additional training. Customization, if needed, happens with end users and experts who provide feedback by testing and using new iterations of a [system] or a pipeline.”

But not every company chooses — or wishes — to go the DIY route. For those preferring a managed solution, there’s the aforementioned Deepset Cloud, which supports customers across the NLP service lifecycle. The service starts with experimentation — i.e., testing and evaluating an app, and adjusting it to a use case, and building a proof of concept — and ends with labeling and monitoring the app in production.

“All NLP services that are developed [with Deepset Cloud] can be used in any end application, simply by integrating an API,” Rusic said. “Example applications are NLP-driven enterprise search (think ‘modern Google-like’ search) and knowledge management.”

With the new financing secured ($15.6 million in total), Deepset aims to translate its open source success — thousands of organizations currently use Haystack — into increased revenue. Rusic says that the 30-person, Berlin, Germany-based company was bootstrapped and break-even before raising its first funding round in 2021, and now has large enterprise customers including Airbus.

“[With the new funding,] we’ll continue to build the open source Haystack NLP project — adding more features, making it even more straightforward for NLP-savvy back-end developers to create NLP services,” Rusic said. “[We’ll also] develop Deepset Cloud into a fully fledged enterprise software-as-a-service to build language-aware applications. This will include enabling more flexible workflows, more granular product lifecycle guidance, and offering essential and supplemental tools, like labeling and data integrations.”

More TechCrunch

It ran 110 minutes, but Google managed to reference AI a whopping 121 times during Google I/O 2024 (by its own count). CEO Sundar Pichai referenced the figure to wrap…

Google mentioned ‘AI’ 120+ times during its I/O keynote

Firebase Genkit is an open source framework that enables developers to quickly build AI into new and existing applications.

Google launches Firebase Genkit, a new open source framework for building AI-powered apps

In the coming months, Google says it will open up the Gemini Nano model to more developers.

Patreon and Grammarly are already experimenting with Gemini Nano, says Google

As part of the update, Reddit also launched a dedicated AMA tab within the web post composer.

Reddit introduces new tools for ‘Ask Me Anything,’ its Q&A feature

Here are quick hits of the biggest news from the keynote as they are announced.

Google I/O 2024: Here’s everything Google just announced

LearnLM is already powering features across Google products, including in YouTube, Google’s Gemini apps, Google Search and Google Classroom.

LearnLM is Google’s new family of AI models for education

The official launch comes almost a year after YouTube began experimenting with AI-generated quizzes on its mobile app. 

Google is bringing AI-generated quizzes to academic videos on YouTube

Around 550 employees across autonomous vehicle company Motional have been laid off, according to information taken from WARN notice filings and sources at the company.  Earlier this week, TechCrunch reported…

Motional cut about 550 employees, around 40%, in recent restructuring, sources say

The keynote kicks off at 10 a.m. PT on Tuesday and will offer glimpses into the latest versions of Android, Wear OS and Android TV.

Google I/O 2024: Watch all of the AI, Android reveals

Google Play has a new discovery feature for apps, new ways to acquire users, updates to Play Points, and other enhancements to developer-facing tools.

Google Play preps a new full-screen app discovery feature and adds more developer tools

Soon, Android users will be able to drag and drop AI-generated images directly into their Gmail, Google Messages and other apps.

Gemini on Android becomes more capable and works with Gmail, Messages, YouTube and more

Veo can capture different visual and cinematic styles, including shots of landscapes and timelapses, and make edits and adjustments to already-generated footage.

Google Veo, a serious swing at AI-generated video, debuts at Google I/O 2024

In addition to the body of the emails themselves, the feature will also be able to analyze attachments, like PDFs.

Gemini comes to Gmail to summarize, draft emails, and more

The summaries are created based on Gemini’s analysis of insights from Google Maps’ community of more than 300 million contributors.

Google is bringing Gemini capabilities to Google Maps Platform

Google says that over 100,000 developers already tried the service.

Project IDX, Google’s next-gen IDE, is now in open beta

The system effectively listens for “conversation patterns commonly associated with scams” in-real time. 

Google will use Gemini to detect scams during calls

The standard Gemma models were only available in 2 billion and 7 billion parameter versions, making this quite a step up.

Google announces Gemma 2, a 27B-parameter version of its open model, launching in June

This is a great example of a company using generative AI to open its software to more users.

Google TalkBack will use Gemini to describe images for blind people

This will enable developers to use the on-device model to power their own AI features.

Google is building its Gemini Nano AI model into Chrome on the desktop

Google’s Circle to Search feature will now be able to solve more complex problems across psychics and math word problems. 

Circle to Search is now a better homework helper

People can now search using a video they upload combined with a text query to get an AI overview of the answers they need.

Google experiments with using video to search, thanks to Gemini AI

A search results page based on generative AI as its ranking mechanism will have wide-reaching consequences for online publishers.

Google will soon start using GenAI to organize some search results pages

Google has built a custom Gemini model for search to combine real-time information, Google’s ranking, long context and multimodal features.

Google is adding more AI to its search results

At its Google I/O developer conference, Google on Tuesday announced the next generation of its Tensor Processing Units (TPU) AI chips.

Google’s next-gen TPUs promise a 4.7x performance boost

Google is upgrading Gemini, its AI-powered chatbot, with features aimed at making the experience more ambient and contextually useful.

Google’s Gemini updates: How Project Astra is powering some of I/O’s big reveals

Veo can generate few-seconds-long 1080p video clips given a text prompt.

Google’s image-generating AI gets an upgrade

At Google I/O, Google announced upgrades to Gemini 1.5 Pro, including a bigger context window. .

Google’s generative AI can now analyze hours of video

The AI upgrade will make finding the right content more intuitive and less of a manual search process.

Google Photos introduces an AI search feature, Ask Photos

Apple released new data about anti-fraud measures related to its operation of the iOS App Store on Tuesday morning, trumpeting a claim that it stopped over $7 billion in “potentially…

Apple touts stopping $1.8B in App Store fraud last year in latest pitch to developers

Online travel agency Expedia is testing an AI assistant that bolsters features like search, itinerary building, trip planning, and real-time travel updates.

Expedia starts testing AI-powered features for search and travel planning