🚨 Soda acquires NannyML! https://launch.soda.io/blog/soda-acquires-nannyml
Soda, leader in data testing and observability, now joins forces with NannyML — open-source library for detecting data drift, concept drift, and silent model failures.
🎯 Why it matters: Soda wants to become the one platform for end-to-end data + ML quality: ✅ Rule-based data checks (pipelines) ✅ ML drift & performance monitoring (production) ✅ No ground truth needed
This move brings Data Engineering and MLOps together on a single platform.
NannyML: https://github.com/NannyML/nannyml
📥 Minimal Permissions for Postgres Ingestion in DataHub
Usually, DataHub needs only metadata, not real data. If you don’t use profiling — reading table rows is not required.
If you want to set permissions safe and precise, here is the minimal working setup:
-- Create user
CREATE USER datahub WITH PASSWORD 'your_strong_password';
-- Access to database and schema
GRANT CONNECT ON DATABASE myproduct TO datahub;
GRANT USAGE ON SCHEMA public TO datahub;
-- Information Schema
GRANT SELECT ON TABLE information_schema.tables TO datahub;
GRANT SELECT ON TABLE information_schema.columns TO datahub;
GRANT SELECT ON TABLE information_schema.views TO datahub;
GRANT SELECT ON TABLE information_schema.schemata TO datahub;
GRANT SELECT ON TABLE information_schema.key_column_usage TO datahub;
GRANT SELECT ON TABLE information_schema.table_constraints TO datahub;
GRANT SELECT ON TABLE information_schema.referential_constraints TO datahub;
GRANT SELECT ON TABLE information_schema.routines TO datahub;
GRANT SELECT ON TABLE information_schema.parameters TO datahub;
-- pg_catalog
GRANT SELECT ON TABLE pg_catalog.pg_class TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_namespace TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_attribute TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_attrdef TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_constraint TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_type TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_enum TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_index TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_inherits TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_depend TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_rewrite TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_proc TO datahub;
GRANT SELECT ON TABLE pg_catalog.pg_roles TO datahub;
-- PostgreSQL function
GRANT EXECUTE ON FUNCTION pg_table_size(regclass) TO datahub;
📌 This gives DataHub access to tables, columns, types, constraints, indexes, and more — only metadata.
Tested and w
🚨 New Post
I just published a deep dive into open-source data quality tools — Great Expectations, Soda, Deequ, DQOps, Spark-Expectations, Pandera, and DQX.
What they’re good at, what they’re missing, and where each one fits best
Read here: Comparative Analysis of Open-Source Data Quality Tools
📦 New version of dag-factory 0.23.0 is out
If you haven’t heard of dag-factory — it’s a tool that lets you create Apache Airflow DAGs from YAML files. No need to write Python code for every DAG. Just define the config and you're done.
In our team, we use dag-factory to generate DAGs that run our Data Quality tests. It works well, configs are simple, and easy to maintain.
I’m also writing this because I contributed to this release:
Updated support for HttpProvider (PR #389)
Added JSON serialization for HttpOperator (PR #382)
Small things, but nice to give back to a tool I use every day.
Tried to connect Trino lineage to DataHub via OpenLineage.
🛠️ Setup: Enabled the OpenLineage Event Listener Sent events to DataHub GMS:
event-listener.name=openlineage
openlineage-event-listener.transport.type=HTTP
openlineage-event-listener.transport.url=https://datahub.internal/openapi/openlineage/api/v1/lineage
openlineage-event-listener.transport.headers=Authorization:Bearer ${ENV:datahub_token}
📉 First attempt failed — DataHub responded with:
java.lang.RuntimeException: Unable to determine orchestrator
🔍 Root cause: DataHub tries to detect the orchestrator based on the producer field using hardcoded logic. Trino sends:
"producer": "https://github.com/trinodb/trino/plugin/trino-openlineage"
…but this isn’t recognized — and DataHub crashes with 500.
👀 There’s a pending fix for this: https://github.com/datahub-project/datahub/pull/13066 It adds an explicit check for Trino:
else if (producer.startsWith("https://github.com/trinodb/trino/")) {
orchestrator = "trino";
}
🧪 Workaround (what worked for me): Just patch the producer field in the JSON to something DataHub understands, like:
"producer": "https://github.com/OpenLineage/OpenLineage/blob/v1-0-0/trino"
Or route events through a small proxy that rewrites it before forwarding.
✅ After that — Trino lineage finally shows up in DataHub.
I spent 8 hours making Trino usage stats work in DataHub. Here's what I learned.
🔧 DataHub ships with starburst-trino-usage — a module designed to pull usage stats (query logs) from a Postgres-based Event Logger.
"You need to setup Event Logger which saves audit logs into a Postgres db and setup this db as a catalog in Trino."
👎 But what if you're not using Starburst?
Trino supports official event listeners:
- HTTP listener
- Kafka listener
- MySQL event listener ✅
- OpenLineage listener
I used the MySQL listener. Sounds simple? Not quite.
🧱 Problem: starburst-trino-usage expects a very specific table schema — but Trino's native listeners log differently.
📄 Solution: You must create a custom VIEW that reshapes raw events into the format starburst-trino-usage expects. Here’s the view I used:
https://gist.github.com/a-chumagin/c9ec0f27d60ca4b10032f0652a1bd034
🪫 It works — but barely. No full control over email formatting, error handling, or scaling across catalogs.
🧠 My recommendation: Use the Trino listener (MySQL, Kafka, etc.), but build your own ingestion in Python with MetadataChangeProposalWrapper. This gives you full control — and works not just for Trino, but for any DB that logs events (even Postgres with pgaudit).
🔥 Bottom line: DataHub gives you tools — but not always flexibility. If you want usage stats done right, especially outside the Starburst bubble — be ready to write code.
📦 I’ll be replacing the module with a custom ingestion script — cleaner, portable, and future-proof.
💡 Why DQ is the next big deal
Why Data Quality will soon become important not only in data world.
GenAI is moving crazy fast — new frontier models are coming almost every month. They get better, faster, cheaper… but there is a limit. Sooner or later, business will start counting money and cut costs. Models will split into two groups:
Work horses — balance between cost and speed
R&D beasts — most expensive and powerful frontier models
And for the “work horses” the key thing will be old principle (GIGO — Garbage In, Garbage Out). If garbage goes in, garbage comes out.
🎯 Example from my practice We have a RAG service (think retrieval-augmented generation) based on Onyx that answers questions on our internal knowledge base.
Problem: a lot of outdated docs, which means the model ends up answering from old or irrelevant data, and half of them not in English, which means more tokens are spent for translation or handling that extra text. No matter how we tune the prompt, it only gets bigger (means more expensive), and keeping it up to date is pain.
Same story with most AI agents. For demo, to impress CEO — perfect. But in production you quickly see: the data they use is the real boss here.
And this is not just my observation — research points exactly the same direction.
📌 Not only us: In Journal of Scientific and Engineering Research link they say:
“AI is determined by the quality of data input that feeds its models… AI models created from low-quality or predominantly biased or incomplete information will produce distortions.” “Data is the lifeblood of any AI model, and the quality of data that feeds into an AI model determines the capability of the resulting model, its precision, stability, and equity.”
And yes, hello from Data Governance — companies want to manage not only quality but also access to data that models can touch.
🎯 Final thought You can replace a model in one day. But cleaning and organizing data — it’s long, boring, and hard work. And it’s the thing that d
Case study: DataHub “Monthly Active Users” ≠ unique people
While digging into the DataHub code, I found that the MAU highlight in the UI does not count unique users.
It counts unique browsers.
- A single person can appear multiple times if they use different devices, incognito, or clear cookies.
- That’s why MAU can be noticeably higher than your actual unique users.
---
How it’s calculated
The GraphQL highlight counts the cardinality of `browserId` over the last month, excluding backend events, using timestamp (event time).
Code proof:
int activeUsersThisRange =
_analyticsService.getHighlights(
_analyticsService.getUsageIndexName(),
Optional.of(dateRangeThis),
ImmutableMap.of(),
ImmutableMap.of(),
Optional.of("browserId"));
public String getUsageIndexName() {
return _indexConvention.getIndexName(DATAHUB_USAGE_EVENT_INDEX);
}
private QueryBuilder getDefaultFilters() {
return QueryBuilders.boolQuery()
.mustNot(
QueryBuilders.termQuery(
DataHubUsageEventConstants.USAGE_SOURCE,
DataHubUsageEventConstants.BACKEND_SOURCE));
}
private AggregationBuilder getFilteredAggregation(
Map<String, List<String>> mustFilters,
Map<String, List<String>> mustNotFilters,
Optional<DateRange> dateRange) {
// Use timestamp as dateRangeField
return getFilteredAggregation(mustFilters, mustNotFilters, dateRange, "timestamp");
}
---
If you want real unique users — query by actorUrn instead of browserId.
All examples below target:
/openapi/v2/analytics/datahub_usage_events/_search
---
1. Unique users (front-end only)
{
"query": {
"bool": {
"must": [
{ "range": { "timestamp": { "gte": "2025-07-13T00:00:00Z", "lt": "2025-08-14T00:00:00Z" } } }
],
"must_not": [
{ "term": { "usageSource": "backend" } }
]
}
},
"aggs": {
"unique_actors": {
"cardinality": {
"field": "actorUrn.keyword",
"precis
🤔 Did you know?
On docs.datahub.com there’s already an Ask AI assistant.
It helps find answers about DataHub, processes, integrations — without scrolling through endless docs.
How it works:
- official docs
- Real Q&A from the Slack community
- DataHub codebase itself
So you can just type in plain English, for example:
- “How do I add a new ingestion source?”
- “Where can I see table lineage?”
…and get an answer right there on the page.
⚠️ Works only in desktop browser version, not mobile.
Handy thing if you didn’t notice it before.🚀
🚨 New Post
I published a post on Habr about how we built a full Data Quality Platform at Ostrovok using open-source tools - Soda Core, Airflow, FastAPI, Streamlit, and DataHub - all running in Kubernetes.
How these tools fit together, why we chose them, and what we learned along the way.
Read here (post in Russian, but easy to translate):
👉 How we built our Data Quality Platform
🍻My second post this week
This is my second story on Habr - about how we built a two-layer metadata architecture at Ostrovok: our own light catalog (DataPortal) + open-source DataHub.
Why keeping both made sense, how we connected them, and what we learned along the way.
Read here (🇷🇺, but translation works fine):
👉 DataHub didn’t replace our self-made data catalog — and that’s okay
From Specification by Example to Spec-Driven Development
It’s not about Data Quality specifically - it’s about QA and software engineering in general.
I’d like to share an interesting direction: Spec-Driven Development (SDD) - a modern take on the classic Specification by Example approach.
I recently tried making changes through a specification- and it worked quite well: the result was a complete specification and Gherkin-based tests.
I planned to use Karate for testing, but since there’s still no proper Python version, I let Claude Code handle the test wrappers.
The Specification as Example approach used to be criticized - engineers had to manually translate Gherkin scenarios into code, which was seen as unnecessary work.
Now, with tools like spec-kit and AI assistants, this problem is gradually disappearing.
It seems that the old idea of specifications by example might be experiencing a quiet renaissance
Understanding Spec-Driven Development
github/spec-kit
Specification by Example
Lately I realized I want to write here not only about data quality, but also about how I use LLMs in my daily work. It’s a big part of my routine now, so it feels natural to share small tips and things that actually help me.
So here’s the first one - a Claude Code pro tip I wish I had learned earlier.
Claude Code Pro Tip - how to reset context properly between tasks
If your project has several connected tasks, don’t wipe the context right away.
It’s much better to save a compact summary first this keeps the important ideas, architecture notes, and working code decisions.
✔️ Recommended flow
1. Compact the current history
Use / compact:
/compact Produce a compact summary focused on:
- key decisions
- relevant code snippets
- file paths
- APIs involved
- next steps
Claude will generate a short structured summary.
2. Save it to your project
Ask Claude to save the summary:
Save the compact summary above into ./memory/task_<id>.md
<id> is any identifier for the current task.
This becomes a portable artifact that won’t disappear.
3. Clear the context (optional)
/ clear
Now you can start the next phase with a clean slate, but without losing knowledge.
4. Load the summary into the next task
At the beginning of the next task:
Here is the summary from the previous task:
@memory/task_<id>.md
Claude gets only the useful information, not 100k tokens of noise.
🎯 Why this matters
• keeps context continuity between tasks
• avoids context bloat and the “dumb zone”
• helps Claude stay in the smart, focused zone
• the compact files become your small “project memory”
• works perfectly long tasks
And yes , add ./memory/task_<id>.md to .gitignore. This is personal working memory, not project code.
Also, a great video on the topic: https://youtu.be/rmvDxxNubIg?si=exdgLQe6owm600wh
🎄 New post on Habr
I finally wrote down something I’ve been thinking about a lot lately - how AI and DG start to merge in real work
The article is about one very concrete thing we built in my company:
DataHub + MCP (Model Context Protocol), so LLMs can actually work with metadata
In short, what the article is about:
• how DataHub turns from a “catalog you click through” into an API for investigations
• how LLMs, via MCP, can answer real questions like impact analysis, lineage, ownership by calling tools
• how this changes everyday work for data engineers, analysts, and data stewards
What I describe is not theory:
• real architecture (Remote MCP, corporate AI chat, ReAct agent)
• real limitations (tokens, batch queries, model choice)
• real examples: “If I delete this table, who will break?”
Why this matters to me personally:
AI only becomes useful in data work when it’s grounded in trusted metadata. And Data Governance suddenly gets a real users who ask questions.
Read here (🇷🇺, auto-translate works fine):
👉 https://habr.com/ru/companies/ostrovok/articles/980210/
In my previous post, I shared how we use MCP with DataHub to make metadata actually usable for Data Practioners.
At the December Town Hall, DataHub announced several features that clearly push in the same direction:
- Ask DataHub - a natural language agent embedded directly into DataHub (cloud-only for now)
- Data Context Graph - a semantic layer connecting structured metadata with unstructured knowledge
- Agent Build Kit - integrations and examples for building agents on top of DataHub as a context platform
I still have questions about how stable and practical this will be in real production setups - the demos had "rough edges".
But the direction makes sense: less clicking, more real questions, grounded in metadata.
Full recap from the DataHub team here:
👉 https://datahub.com/blog/datahub-town-hall-building-ai-agents/
Claude Code as Debug tool
Claude Code isn’t just about “writing a service” or “creating an app.” It’s also very well suited for debugging.
Recently at work, I had to troubleshoot a problem between Metabase, DataHub, and the database. The metadata was inconsistent, and I needed to understand where and why this was happening.
What I needed to do:
- extract SQL query metadata from Metabase and DataHub
- parse the SQL and extract the tables
- compare them with each other
- verify that these tables actually exist in the database
- figure out at which point everything was breaking
I used Claude Code with the DataHub MCP enabled. I described the problem and said upfront:
"You can write any scripts in the sandbox and connect to Metabase, DataHub, and the database."
The result was an almost fully automated investigation:
the agent accessed the systems itself
- collected and compared metadata using scripts (the MCP was ultimately not needed)
- looked into the application code where the problem was
- found the root cause
- discovered several additional issues no one had suspected
The conclusion is simple: it’s possible and necessary to automate not only application creation, but also debugging, investigations, and incident reviews.
Have you tried using Claude Code for Data Quality?
⚠️ Heads-up for DataHub users with custom lineage ingest
If you have custom lineage ingest (for example via OpenLineage), be careful with Tableau metadata ingestion.
Problem
- Custom lineage shows correctly at first.
- After running Tableau ingestion, lineage disappears in UI and GraphQL.
- There is currently no config to disable or protect existing lineage from Tableau ingestion.
Why this happens
Tableau ingestion generates different upstream URNs. It strips the schema when the table name already contains it.
Example:
- Expected: urn:li:dataset:(dataPlatform:iceberg,db.schema.table,PROD)
- Generated by Tableau ingest: urn:li:dataset:(dataPlatform:iceberg,db.table,PROD)
These schema-less datasets don’t exist in the catalog => graph lineage becomes empty.
Symptoms
- upstreamLineage aspect exists, but points to non-existing datasets
- dataset.lineage (GraphQL) returns empty
- UI lineage is empty
- Reindex does not help
Current workaround
- Re-run custom lineage ingest after every Tableau ingestion
- Yes, it’s painful.
I’ve opened an issue in DataHub with full details:
https://github.com/datahub-project/datahub/issues/16018
If you rely on custom SQL lineage for Tableau and probably with Metbase - keep this in mind.
In the last two weeks, I was deeply involved in data migration testing. I shared some thoughts about migration testing in one of my previous posts
But today I want to talk about one mistake I see again and again in migration pipelines.
The main problem in data migration is the temptation to run full checks immediately. A full check is usually a cross join comparison between source and target tables. It sounds powerful. It feels serious. But in reality, it is the most expensive test in the whole pipeline.
And very often it fails not because of some deep data issue, but because of simple things - row count mismatch or column inconsistency.
My advice for organizing a migration testing pipeline is simple: start from cheap and fast tests (like smoke tests in SQA). Compare row counts. Check column consistency. Make sure schemas match.
The next step is profiling. Compare aggregations between two tables. If you have business checks — run them here. This stage already gives a lot of confidence.
Only after that move to full cross join comparison. These tests should run only when line checks and profiling/business checks have passed.
If a full check fails, it often means you missed something earlier. Improve profiling. Shift more checks to the left.
Of course, do not forget about security and performance testing.
A structured test suite saves time and energy. Trying to validate everything with one “super test” usually creates more noise than value.
Move step by step.