Slack Fast-Tracks Metadata Visibility with DataHub

Slack Case Study

CUSTOMER
Slack
INDUSTRY
Technology
SIZE
3,500+ employees
SOLUTION
DataHub Core (OSS)
USE CASE
Discovery, Lineage
DATA STACK
AWS, Hive Metastore, Airflow, Presto, Spark, Kafka, Kinesis, MySQL, Thrift, Docker

GOALS

The Topline

Note: This story was originally published July 2023.

Challenge

Slack’s data engineering team faced a persistent six-year challenge: building a unified metadata layer across a highly complex ecosystem. Despite having what Senior Data Engineer, Nedra Albrecht, described as a “very best practice typical warehouse construction” with AWS, Hive Metastore, Airflow, Presto, and Spark, crucial data context remained elusive.

“Where I think things get complicated is that there’s this data context layer on top of [the stack]… we don’t often see that. If we just get our system architecture right, it’s just gonna magically come together. But really, there’s a lot of complexity.”
Nedra Albrecht, Senior Data Engineer, Slack

Key pain points included:

Over the years, Slack trialed multiple solutions, from Airflow-based dashboards and internal tools to OpenSearch, Marquez, and ANTLR query parsers, each falling short.

“We’ve tried a lot. And we learned a lot of things,” Albrecht reflected. Each attempt taught them that “managing metadata is not only a complex process, it requires flexibility in thinking about it.”

Solution

Slack found its breakthrough with DataHub. What stood out immediately was DataHub’s support for both push and pull metadata workflows, eliminating the need to choose between them as previous attempts had required.

Other standout capabilities included:

“All I have to do is inject the data as we have it either through recipes that already exist or my own custom data that I’m injecting. We can set up governance policies to help ease the burden on data engineering, empower our users, really do important things for them.”
Nedra Albrecht, Senior Data Engineer, Slack

The team deployed DataHub in just three days as part of a Hack Day initiative, using Docker containers and Slack’s internal orchestration tool, Bedrock. “It was really, really easy, actually,” emphasized Albrecht.

Impact

Slack saw immediate, measurable results that transformed how metadata is managed across the organization.

Key outcomes included:

“This tool is so flexible. The data model itself is extensible, which I love. It’s based on a graph network, which is the way metadata should be represented. It is, to me, just mind-blowingly good.”
Nedra Albrecht, Senior Data Engineer, Slack

DataHub transforms enterprise metadata management with AI-powered discovery, intelligent observability, and automated governance.