Automated Data Warehouse Architecture to Support Knowledge Management in Patent Databases

Autores

DOI:

https://doi.org/10.5212/31xd4e05

Resumo

This work presents an automated Data Warehouse architecture designed to support the structuring, analysis, and management of knowledge in patent databases. Based on XML files provided by the United States Patent and Trademark Office (USPTO), a Python-based Extract, Transform, Load (ETL) pipeline was developed to preprocess, standardize, and load data into a PostgreSQL database modeled using a star schema. The architecture, executed through Docker containers, integrates a staging layer, analytical dimensions, a fact table, a REST API, and OLAP queries. In the initial load, 189,410 raw records were processed and consolidated into 16,311 unique patents and 651,117 analytical associations. Textual processing identified 25,658 distinct terms through tokenization, while Artificial Intelligence-based semantic enrichment added 45,449 technical terms. The analyses supported the exploration of term frequencies, temporal evolution, technological categories, thematic clusters, semantic co-occurrences, and trend forecasts. The results demonstrate that the multidimensional model enables temporal, semantic, technological, geographical, and authorship perspectives to be combined within an integrated structure. Therefore, the proposed solution provides a reproducible and extensible analytical foundation for supporting technology forecasting, trend monitoring, and decision-making in research, development, and innovation.

Downloads

Publicado

2026-09-25