LEGAT – Advanced computer system based on artificial intelligence (AI) for identifying and extracting entities from unstructured data collections
News
2026
-
[Jun'26] -- Milestone reached: Execution closing report.
-
[May'26]
-
Organized dissemination session with participants from the public, private and academic sectors;
-
Milestone reached: Fully tested and demonstrated prototype system under
conditions relevant to the beneficiary (TRL 8);
-
[Apr'26] -- Milestone reached: Prototype demonstrated functionally under
relevant operating conditions (TRL7);
-
[Mar'26] -- Invited the beneficiary for prototype testing;
-
[Feb'26] -- Milestone reached: Integrated preliminary product validated under relevant operating conditions (TRL5);
-
[Jan'26] -- Started final internal testing round;
2025
-
[Dec'25] -- Installation is live at the UB datacenter.
-
[Nov'25]
-
Finilized the acquisition process for the equipments;
-
Started the testing and functional validation round;
-
[Oct'25] -- Milestone reached: Integrated preliminary product (hardware-software
platform) validated in laboratory environment (TRL4);
-
[Sep'25] -- Integration process raised a few difficult issues, fixing;
-
[Jun'25] -- Milestone reached: Solution validated in laboratory environment (TRL4);
-
[May'24] -- AI, document extraction, UI, database and search modules implemented;
-
[Feb'24] -- Started the aquisition process for the equipments;
-
[Jan'24] -- Started implementation of AI, document extraction, UI, database and
search modules;
2024
-
[Dec'24]
-
Milestone reached: Technology concept formulated (TRL2);
-
We have a new member, welcome Adriana Costache!;
-
[Nov'24] -- Prepared an initial demo for the beneficiary;
-
[Oct'24] -- Started hiring process for one PhD position;
-
[Sep'24] -- Milestone reached: state of the art and existing solutions report;
-
[Jun'24] -- Kickoff!
Project
-
Project ID: PN-IV-P6-6.3-SOL-2024-0090
-
Consortium: UB (coordinator),
ATM,
Nextgen Software SRL.
-
Team: 30 positions
-
Funder:
UEFISCDI
-
Budget: 2.122.787 lei (~ 424.557 euro)
-
Duration: 06 June 2024 - 05 June 2026
Main Objective
The main objective of this project is
the creation of a hardware-software IT system, called LEGAT,
based on artificial intelligence, which,
based on certain training data sets,
will proceed to the semi-automatic structuring of the historical data
collected at the MAI/DGPI level
based on on a number of essential components:
-
extracting data from unstructured datasets;
-
characterization of the entities and the links between them;
-
identifying patterns and retrieving the information of the entities in focus.
End Result
Prototype hardware-software platform delivered to the Beneficiary at the end of the project.
Team
Paul Irofti -- Project Coordinator
University of Bucharest:
Paul Irofti -- Principal Investigator
Radu Ionescu -- Senior Researcher
Marius Popescu -- Senior Researcher
Iulia Timofte -- Researcher
Roxana Voicu -- Researcher
Eduard Poesina -- Assistant Researcher
Silviu Gheorghe -- Master Student
Adriana Costache -- Assitant Researcher (started December 2024)
Ana Cristina Rogoz -- Assistant Researcher (withdrawn before kickoff)
Military Technical Academy
Luciana Morogan -- Principal Investigator
Ion Bica -- Senior Researcher
Ștefan-Adrian Toma -- Senior Researcher
Mihai Coca -- Researcher
Iulian Tiță -- Assistant Researcher
Mirabela Medvei -- Assistant Researcher
George Hariga -- Assistant Researcher
Alexandra Buzățoiu -- Master Student
Paul-Florinel Căsăndroiu -- Master Student
Ilie-Cosmin Bilțan -- Master Student
Florina Conchințoiu -- L1 Technician
Andrei Brînzea -- L1 Technician
Nextgen Software SRL:
Bogdan Legănaru -- Principal Investigator
Vlad Gladin -- Senior Researcher
Daniel Tache -- Researcher
Alexandru Cocosila -- Researcher
Viorel Tiganescu -- L2 Technology Engineer
Bonciu Emilian Cristian -- L2 Technology Engineer
Adrian Bogdan Sandu -- L2 Technician
Documentation
Papers
|
[1]
|
M. Marogel,
and
S.F. Gheorghe,
“Munibuc at Touché: Generalist Embeddings for Ideology and Populism Detection,”
in CLEF 2025 Conference and Labs of the Evaluation Forum Touché:
Argumentation Systems, 2025, pp. 1--6.
[ PDF |
http ]
|
|
[2]
|
A.V. Costache,
S.F. Gheorghe,
E.G. Poesina,
P. Irofti,
and
R.T. Ionescu,
“Text Classification Under Class Distribution Shift: A Survey,”
in Proceedings of the 19th Conference of the European Chapter of
the Association for Computational Linguistics (EACL), 2026, pp. 4045--4060.
[ PDF |
DOI |
http ]
|
|
[3]
|
A.C. Rogoz,
R.T. Ionescu,
A.V. Anghel,
I.L. Antone-Iordache,
S. Coniac,
and
A.I. Ionescu,
“A large-scale benchmark for evaluating large language models
on medical question answering in Romanian,”
npj Digital Medicine, vol. 9, pp. 268, 2026.
[ PDF |
DOI |
http ]
|
|
[4]
|
A.V. Costache,
D.N. Dragomir,
S.V. Gheorghe,
E. Poesina,
P. Irofti,
and
R.T. Ionescu,
“MOSLD-Bench: Multilingual Open-Set Learning and Discovery Benchmark
for Text Categorization,”
pp. 1--9, 2026 (under review).
[ PDF |
arXiv ]
|
|
[5]
|
A.V. Costache,
E. Poesina,
S.V. Gheorghe,
P. Irofti,
and
R.T. Ionescu,
“Multilingual Coreference Resolution via Cycle-Consistent Machine Translation,”
pp. 1--9, 2026 (under review).
[ PDF |
arXiv ]
|
|
[6]
|
G.A Dima
I.C. Biltan
M Medvei
and
L.M Morogan
“RoRED: A Romanian Relation Extraction Dataset,”
AI, vol. 7, no. 6, pp. 225, 2026.
[ PDF |
DOI |
http ]
|
|
[7]
|
A Gavrilaș
and
L.M Morogan
“Latency, Reasoning, and Hallucination:
A Logical DoS Benchmark for Open-Source LLMs,”
in The 16th International Conference on Communications,
2026, pp. 1--6.
[ PDF ]
|
About
LEGAT aims to create a hardware-software computer system based on artificial intelligence, which, based on training data sets, will proceed to the semi-automatic structuring of the historical data collected at the MAI/DGPI level based on a series of essential components:
(i) extracting data from unstructured data sets
Training large language models (LLM) for a high degree of accuracy and efficiency starting from pre-trained models, which we will adapt to our data sets, using effective training techniques such as Low-Rank Adaptation (LoRA), Direct Preference Optimization (DPO) or combinations thereof. To maximize performance, we will manually annotate the data and fit LLM models in a supervised manner.
(ii) characterization of the entities and the links between them;
Starting from the latent representation obtained by the language model trained by our team, we will add a module consisting of neural layers for entity extraction that will have a classification layer that assigns to each language token a class representing an entity type or a class representing simple words (non-entities). We will use a module with a similar architecture for identifying and finding attributes for entities. A third neural module will be used to extract relationships.
(iii) identifying patterns and retrieving the information of the entities in focus
The resulting data will be inserted into dedicated database tables. Also, the metadata of the processed file will be inserted as well as other information considered necessary for the creation and identification of predefined topics. The web interface module will value the entities presented in the database with multiple query criteria: links or property values. Thus, users can investigate entities, links between entities and a map of links between them.