
Our product is a data platform that bundles open source projects such as Apache Kafka, Apache Airflow, Apache Spark and Trino, each in several supported versions. Because of that vast scope, a single platform release contains over 50,000 reported vulnerabilities across its roughly 40 product and operator images. With two supported releases at any time we face that scope twice over, which is far more than our security team can analyze by hand.
This talk describes our journey from being swamped to running an efficient, largely automated vulnerability handling process. We start with the foundations: building every product from source and generating accurate SBOMs, without which everything downstream is guesswork. We then show how we cut through the noise by encoding triage logic as Rego policy rules over metrics like EPSS, CVSS and KEV catalogs in the open source tool SecObserve. Finally, we present our AI-assisted assessment flow: one LLM performs the exploitability analysis with access to source code and platform deployment context, a second independent model reviews it, and a human makes the final call before results are published as VEX documents. Roughly 300 in-depth vulnerability analyses so far covered about 14,000 individual occurrences, humans rejected under one percent of the AI verdicts, and our per-vulnerability effort dropped substantially without sacrificing quality. We share what we learned and improved along the way, what human review still catches, and our honest take on when, if ever, the human can step out of the loop.

Conference partners







Organiser
