🌸 Spring Data Allergy: An Evening at SumUp That I Did Not Walk Away From Unaffected
Let me be upfront: the meetup was brilliant. The talks were sharp, the energy in the room was electric, and the conversations ran well. The only thing I did not manage to protect myself from, despite my best efforts, was catching a full-blown case of data enthusiasm (an occupational hazard I have clearly not built immunity to yet). Consider this my delayed sneezing fit, in blog form.
Setting the Scene: A Spring Evening at Koppenstraße 8
On the evening of April 23rd, a few dozen data professionals made their way to the SumUp offices in Friedrichshain, tucked into the kind of Berlin neighbourhood that still has enough character to remind you this city runs on more than venture capital and oat milk.
Data Berlin is a community gathering of professionals and enthusiasts working in data analytics, business intelligence, and data engineering Meetup and this edition, titled Spring Data Allergy, leaned fully into the seasonal metaphor. Spring is here. And so are the sniffles. Not from pollen, but from data pipelines, AI hype, and the occasional rogue SQLs; from a company that, incidentally, has an enormous amount of real-world data coursing through its systems daily. The setting was fitting.
The Talks: Where the Allergy Really Kicked In
Three talks, three teams, one recurring theme: the gap between what AI promises and what your data actually supports.
Johana Yuan Feng opened with something I find genuinely exciting using multi-LLM orchestration to turn dense regulatory language into automated compliance pipelines. The kind of work that makes you think: if we can parse a regulatory document and build a SQL pipeline from it, what else is just a translation problem waiting to be solved?
Marina Ariamnova then delivered the honest talk of an AI had been generating analytics outputs for three months. No one noticed they were fabricated. The root cause was not the model, it was that only a narrow slice of their data was actually AI-ready, and the system had no mechanism to say so. The lesson is one the field keeps needing to relearn, you cannot skip the data foundation and go straight to the intelligence layer. The research on this is unambiguous: poor data quality costs organisations not just money but trust, and once an AI erodes that trust, rebuilding it takes considerably longer than the three months it took to lose it (Redman, 2016).
Dafne Romero closed with a framework her team, built where LLMs act as judges over real customer support conversations that scoring resolution quality, flagging PII leakage, evaluating human agents across channels, and classifying formal complaints for compliance purposes. What made it compelling was not the architecture, but the honesty about what makes it hard: prompts drift, transcripts are inconsistent, models are non-deterministic, and every stakeholder has a different definition of "good." That tension, between what evaluation frameworks promise in theory and what they require to survive in production, is exactly where the most important work in applied AI is happening right now.
Taken together, the three talks made a quiet argument: data is not the new oil. Reliable, well-governed, honestly evaluated data is. The rest is just pressure without a pipeline.
What Data Can Already Do — and What We Are Still Leaving on the Table
Here is where I have to be honest about the thing that hit me hardest that evening, and it was not any single slide. It was the accumulated weight of realising how much is already possible and how large the gap remains between what the field knows and what most organisations actually practise.
We have techniques today that allow us to detect anomalies in real-time transactional data before a human analyst would notice anything is wrong. We can model customer churn with enough lead time to act. We can surface the KPIs that correlate with revenue in ways that challenge every assumption a leadership team has held for years. We can automate the reporting layer entirely and redirect analyst capacity toward interpretation, strategy, and the uncomfortable questions that numbers alone cannot answer.
And yet, in most organisations I have worked with or spoken to the data team is still producing weekly Excel exports for people who have already made up their minds.
This is not a technology gap. The tools exist. The research exists. The frameworks exist. It is a translation gap: the distance between what data can tell you and what decision-makers are equipped to hear. Closing that gap is, I would argue, the most important work in this field right now.
The research broadly confirms this tension. A landmark study published in the MIT Sloan Management Review by McAfee and Brynjolfsson established that data-driven organisations outperform peers across productivity and profitability metrics and that evidence predates the current wave of AI capability by over a decade. We have known this. The question has never been whether data works. It is why adoption is still so uneven.
Why I Keep Showing Up to These Things
I will be direct: I do not attend meetups for the free drinks, though I appreciate them. I attend because there is something that happens in a room of people who are all genuinely trying to figure out the same hard problems that does not happen anywhere else. You hear someone describe a solution you have been circling for months. You realise your frustration is shared. You leave with three things to read, one question you had not thought to ask, and a renewed sense that the work matters.
Spring Data Allergy delivered on all three counts. The setting was generous, the talks were substantive, and the conversations in the networking stretches were the kind that run over time for good reasons.
The next edition will find me there. Possibly with antihistamines. Definitely without regrets.
Did you attend Spring Data Allergy? I would love to hear what stayed with you. Find me on LinkedIn or drop a comment below.
References
- Biskupski, T. & Kleber, S. (2025). Evaluating the reliability and fidelity of automated judgment systems of large language models. arXiv preprint. Available at: https://arxiv.org/pdf/2603.22214 [Accessed 27 April 2026].
- Gao, J., Chen, C., Jia, Y., Gong, X., Lam, K. & Wang, Q. (2025). Evaluating and mitigating LLM-as-a-judge bias in communication systems. arXiv preprint. Available at: https://arxiv.org/pdf/2510.12462 [Accessed 27 April 2026].
- Gu, J. et al. (2024). LLMs-as-judges: A comprehensive survey on LLM-based evaluation methods. arXiv preprint arXiv:2412.05579. Available at: https://arxiv.org/abs/2412.05579 [Accessed 27 April 2026].
- Redman, T.C. (2016). Bad data costs the U.S. $3 trillion per year. Harvard Business Review, 22 September. Available at: https://community.sap.com/t5/technology-blog-posts-by-sap/bad-data-costs-the-u-s-3-trillion-per-year/ba-p/13575387 [Accessed 27 April 2026].
- Jiang, S. et al. (2025). 'Stars at the Regulations Challenge Task: A large language model for financial regulation', in Proceedings of the Joint Workshop of the 9th FinNLP, the 6th FNP, and the 1st LLMFinLegal, pp. 371–384. Available at: https://aclanthology.org/2025.finnlp-1.43.pdf [Accessed 27 April 2026].
- Responsible Innovation Framework (2025). Responsible innovation: A strategic framework for financial LLM integration. arXiv preprint arXiv:2504.02165. Available at: https://arxiv.org/html/2504.02165v1 [Accessed 27 April 2026].