Enterprise

Is the modern data stack just old wine in a new bottle?

Comment

Bottle in a paper bag on a gray background. Dark bottle of alcohol in a crumpled brown bag. Close-up. Selective focus.
Image Credits: Mikhail Dmitriev (opens in a new window) / Getty Images

Ashish Kakran

Contributor
Ashish Kakran, principal at Thomvest Ventures, is a product manager/engineer turned investor who enjoys supporting founders with a balance of technical know-how, customer insights, empathy with challenges and market knowledge.

More posts from Ashish Kakran

Remember the cable, phone and internet combo offers that used to land in our mailboxes? These offers were highly optimized for conversion, and the type of offer and the monthly price could vary significantly between two neighboring houses or even between condos in the same building.

I know this because I used to be a data engineer and built extract-transform-load (ETL) data pipelines for this type of offer optimization. Part of my job involved unpacking encrypted data feeds, removing rows or columns that had missing data, and mapping the fields to our internal data models. Our statistics team then used the clean, updated data to model the best offer for each household.

That was almost a decade ago. If you take that process and run it on steroids for 100x larger datasets today, you’ll get to the scale that midsized and large organizations are dealing with today.

For example, a single video conferencing call can generate logs that require hundreds of storage tables. Cloud has fundamentally changed the way business is done because of the unlimited storage and scalable compute resources you can get at an affordable price.

To put it simply, this is the difference between old and modern stacks:

Image Credits: Ashish Kakran, Thomvest Ventures

Why do data leaders today care about the modern data stack?

Self-service analytics

Citizen-developers want access to critical business dashboards in real time. They want automatically updating dashboards built on top of their operational and customer data.

For example, the product team can use real-time product usage and customer renewal data for decision-making. Cloud makes data truly accessible to everyone, but there is a need for self-service analytics compared to legacy, static, on-demand reports and dashboards.

Serving predictions

Once machine learning models are trained and ready to be used, there needs to be an easy way for different teams within an organization to benefit from them. This is typically achieved via a simple URL that accepts requests and returns predictions. Building these microservices and maintaining them is a core challenge when you are serving thousands of HTTP requests per second.

Data transformation

Data scientists want to be able to track older versions of data so that they can run experiments and know what version of data was used to complete training. This need is creating popular products that are optimized for in-place transformation of data.

Data quality

Some cutting-edge data organizations now prefer a data-centric approach to a model-centric approach. The belief that more data means better results is being replaced by the belief that the quality of data matters more. Typically, trained models are observed using two parameters, precision and recall. Precision tells you the proportion of positive identification that was actually correct, and recall tells you the proportion of actual positives that were correctly identified. Now imagine ensuring data quality for real-time data streams coming at you in a variety of different formats.

How do the legacy and modern data stacks compare?

Generally speaking, the modern data stack is about leveraging cloud resources to more effectively analyze complex streaming data.

Image Credits: Ashish Kakran, Thomvest Ventures

Here are a few key trends that enterprises should note:

  • The ETL process is becoming EL (T), which means the data is first dumped as it is received in certain locations like a data lake. This way, the storage systems don’t complain about the format of data as it is stored. Once the data is stored, then it can be processed in-place for analytics. By doing this, the firehose of continuous data can be more effectively managed, processed and analyzed.
  • Data observability has become critical. Data fails silently, and with rapidly evolving data stacks, it is necessary to be able to monitor data and set alerts to fix issues. You don’t want your trained models that teach Spanish to accidently train on English words or on missing data. One just can’t visually analyze and fix millions of rows of data.
  • The emergence of the chief data/AI/data and analytics officer. Data is such a complex problem that CIOs now have CDOs/CAOs/CDAOs reporting to them. While we started the 21st Century talking about data as competitive advantage, we are now in a time when unmanaged data becomes toxic. There are regulatory laws about how data can be used, shared or handled. How do you comply with a customer’s request to delete all their data if you don’t even know where and in what form it is stored in?

Opportunities

Each step of the data analysis process is ripe for disruption. While visionary founders are building cloud-native tools to win emerging data categories, the incumbents have been slower to react. Whether building data pipelines or ML pipelines, organizations today have a variety of open and closed source technologies to choose from.

Image Credits: Ashish Kakran, Thomvest Ventures

Practitioners are spoilt for choices when building enterprise data pipelines.

Image Credits: Ashish Kakran, Thomvest Ventures

The efficient data stack for data engineers, database developers and data scientists changes every four to five years. Companies moved to big data analytics to analyze large datasets in private data centers, and though it promised many benefits, big data remains technically complex to implement. The modern data stack makes this easy by leveraging the scale, reliability and resilience of the cloud.

The rules are being rewritten on how data will be used for competitive advantage, and it won’t be long before the winners emerge. Incumbents are redesigning their legacy software to run on the cloud, but our bet is on nimble teams run by visionary founders.

More TechCrunch

South Korea’s fabless AI chip industry saw a slew of fundraising events over the last couple of years as demand for hardware to power AI applications skyrocketed, and it seems…

Fabless AI chip makers Rebellions and Sapeon to merge as competition heats up in global AI hardware industry

Here’s a list of third-party apps that were Sherlocked by Apple at this year’s WWDC.

The apps that Apple Sherlocked at WWDC 2024

Black Semiconductor, which is developing a chip-connecting technology based on graphene, has raised $273M in a combination of private and public funding. 

Germany’s Black Semiconductor raises $273M for graphene-based chip connectivity tech

Featured Article

Let there be Light! Danish startup exits stealth with $13M seed funding to bring AI to general ledgers

It’s not the sexiest of subject matters, but someone needs to talk about it: The CFO tech stack — software used by the chief financial officers of the world — is ripe for disruption. That’s according to Jonathan Sanders, CEO and co-founder of fledgling Danish startup Light, which exits stealth…

3 hours ago
Let there be Light! Danish startup exits stealth with $13M seed funding to bring AI to general ledgers

Fresh off the success of its first mission, satellite manufacturer Apex has closed $95 million in new capital to scale its operations.  The Los Angeles-based startup successfully launched and commissioned…

Apex’s off-the-shelf satellite bus business attracts $95M in new funding

After educating the D.C. market, YC aims to leverage its influence, particularly in areas like competition policy.

DC’s political class doesn’t know Y Combinator exists — yet

Lina Khan says the FTC wants to be effective in its enforcement strategy, which is why it has been taking on lawsuits that “go up against some of the big…

FTC Chair Lina Khan tells TechCrunch the agency is pursuing the ‘mob bosses’ in Big Tech

With dozens of antitrust cases and close to a hundred on the consumer protection side, the agency is now turning to innovative tactics to help it fight fraud, particularly in…

FTC Chair Lina Khan shares how the agency is looking at AI

The ability to pause your activity rings is a minor feature update for most, but for those of us who obsess about such things to an unhealthy degree, it’s the…

Apple Watch is finally adding a feature I’ve been requesting for years

Featured Article

Why Apple is taking a small-model approach to generative AI

It’s a very Apple approach in the sense that it prioritizes a frictionless user experience above all.

11 hours ago
Why Apple is taking a small-model approach to generative AI

When generative AI tools started making waves in late 2022 after the launch of ChatGPT, the finance industry was one of the first to recognize these tools’ potential for speeding…

Linq raises $6.6M to use AI to make research easier for financial analysts

In addition to the federal funding, the state of New Mexico — where SolAero is based — committed to providing financing and incentives that value $25.5 million.

Biden administration looks to give Rocket Lab $24M to boost space-grade solar cell production

Some of the new Apple Intelligence features that Apple debuted at WWDC 2024 don’t even feel like AI, they just feel like smarter tools. 

Apple’s AI, Apple Intelligence, is boring and practical — that’s why it works

The TechCrunch team runs down all of the biggest news from the Apple WWDC 2024 keynote in an easy-to-skim digest.

Here’s everything Apple announced at the WWDC 2024 keynote, including Apple Intelligence, Siri makeover

Jordan Meyer and Mathew Dryhurst founded Spawning AI to create tools that help artists exert more control over how their works are used online. Their latest project, called Source.Plus, is…

Spawning wants to build more ethical AI training datasets

After leading the social media landscape, TikTok appears to be interested in challenging Google’s dominance in search. The company confirmed to TechCrunch that it’s testing the ability for users to…

TikTok comes for Google as it quietly rolls out image search capabilities in TikTok Shop

General Motors is investing $850 million into Cruise as the autonomous vehicle subsidiary slowly makes its way back to testing in Phoenix, Dallas and, as of Tuesday, Houston. GM’s CFO…

GM gives Cruise $850M lifeline as it relaunches robotaxis in Houston

These messaging features, announced at WWDC 2024, will have a significant impact on how people communicate every day.

At last, Apple’s Messages app will support RCS and scheduling texts

Welcome to TechCrunch Fintech! This week, we’re looking at Rippling’s controversial decision to ban some former employees from selling their stock, Carta’s massive valuation drop, a GenZ-focused fintech raise, and…

Rippling’s tender offer decision draws mixed — and strong — reactions

Google is finally making its Gemini Nano AI model available to Pixel 8 and 8a users after teasing it in March.

Google’s June Pixel feature drop brings Gemini Nano AI model to Pixel 8 and 8a users

At WWDC 2024, Apple introduced new options for developers to promote their apps and earn more from them in the App Store.

Apple adds win-back subscription offers and improved search suggestions to the App Store

iOS 18 will be available in the fall as a free software update.

Here are all the devices compatible with iOS 18

The acquisition comes as BeReal was struggling to grow its user base and was looking for a buyer.

BeReal is being acquired by mobile apps and games company Voodoo for €500M

Unlike Light’s older phones, the Light III sports a larger OLED display and an NFC chip to make way for future payment tools, as well as a camera.

Light introduces its latest minimalist phone, now with an OLED screen but still no addictive apps

Since April, a hacker with a history of selling stolen data has claimed a data breach of billions of records — impacting at least 300 million people — from a…

The mystery of an alleged data broker’s data breach

Diversity Spotlight is a feature on Crunchbase that lets companies add tags to their profiles to label themselves.

Crunchbase expands its diversity-tracking feature to Europe

Thanks to Apple’s newfound — and heavy — investment in generative AI tech, the company had loads to showcase on the AI front, from an upgraded Siri to AI-generated emoji.

The top AI features Apple announced at WWDC 2024

A Finnish startup called Flow Computing is making one of the wildest claims ever heard in silicon engineering: by adding its proprietary companion chip, any CPU can instantly double its…

Flow claims it can 100x any CPU’s power with its companion chip and some elbow grease

Five years ago, Day One Ventures had $11 million under management, and Bucher and her team have grown that to just over $450 million.

The VC queen of portfolio PR, Masha Bucher, has raised her largest fund yet: $150M

Particle announced it has partnered with news organization Reuters to collaborate on new business models and experiments in monetization.

AI news reader Particle adds publishing partners and $10.9M in new funding