AI-Powered Topic Discovery and Text Quality Analytics
- Project
- 22017 CAPE
- Type
- New product
- Description
The exploitable result is a scalable and quality-controlled NLP analytics solution for analysing large volumes of passenger feedback and other domain-specific textual data. It addresses the difficulty organisations face in extracting reliable insights from unstructured comments when real-world labelled data are limited, imbalanced or insufficient for training and validating AI models. Unlike generic public datasets, the solution generates scalable and balanced textual datasets according to predefined categories and scenarios, reducing dependency on scarce or poorly structured real-world data. Topic assignments are evaluated at individual-comment level using purity scores, rather than relying only on conventional cluster-level coherence measures. This provides more granular evidence of whether identified topics genuinely represent the underlying comments.
- Contact
- Aylin Yorulmaz
- aylin.yorulmaz@kocsistem.com.tr
- Research area(s)
- Natural Language Processing (NLP), Data Centric AI, Explainable AI (XAI), Unsupervised Learning
- Technical features
1- Large-scale dataset processing: Supports the expansion and analysis of textual datasets from hundreds towards 100,000 passenger comments for more reliable model training and evaluation.
2- Embedding-based text representation: Transforms comments into semantic vector representations using sentence embeddings, enabling meaning-based comparison and clustering.
3- BERTopic-based topic modelling: Combines UMAP dimensionality reduction, HDBSCAN density-based clustering and c-TF-IDF topic extraction to identify themes within unstructured text.
4 - Comment-level topic validation: Calculates purity scores for individual topic assignments and analyses topic-to-category correspondence, providing more granular validation than cluster-level assessment alone.
- Integration constraints
The analytical pipeline can be exposed through REST endpoints for submitting comments and retrieving predicted topics, category mappings, purity scores, outlier indicators and quality metrics. The individual processing modules can be embedded in an existing data-science or AI pipeline.
- Targeted customer(s)
Airports, airlines, public transport operators, mobility service providers, customer-experience management companies, and organisations analysing large volumes of customer feedback.
- Conditions for reuse
The exploitable result is intended to be reused under a commercial licensing model, either as a standalone NLP analytics component or as part of a broader customer-experience, business-intelligence or data-analytics solution. The core know-how, algorithms, evaluation methodologies, synthetic data generation approaches and implementation artefacts remain the intellectual property of the project partner(s) developing the solution.
- Confidentiality
- Public
- Publication date
- 25-09-2026
- Involved partners
- KoçSistem (TUR)
- inosens (TUR)