Mykolas Romeris University Research Management System (CRIS)





Use this url to cite department: https://cris.mruni.eu/cris/handle/007/20066
Now showing 1 - 10 of 11
  • dataset[2020][D][H004];
    Vytauto Didžiojo universitetas, 2020

    The corpus is comprised of 154 EU legislative documents (English documents and their translations into French and Lithuanian) related to various financial issues and enacted in the period 2013-2018. The documents were extracted from the Official Journal of the European Union available in the open EUR-Lex database. All documents were converted into plain text format and sentence-aligned. The size of the corpora is as following: EN 1 006 485 words, FR 1 181 647 words, LT 803 845 words.

      26
  • dataset[2022][D][H004]
    Utka, Andrius
    ;
    ; ; ; ;
    Vytauto Didžiojo universitetas, 2022

    The English-Lithuanian comparable corpus (DVITAS COMPARABLE) is morphologically annotated. It includes English and Lithuanian original texts on cybersecurity from the time period of 2010-2021. The corpus was compiled for the bilingual terminology extraction project together with English-Lithuanian parallel corpus. There are 1,708 files in English and 2,567 for Lithuanian. The total size of the corpus is 4m words (EN-2m; LT-2m) The corpus is composed of texts representing 4 text types: academic (EN-19%; LT-30%), administrative-informative (EN-8%; LT-11%), legal (EN-18%; LT-4%), media (EN-55%; LT-55%).

      60
  • dataset[2022][D][H004];
    Vytauto Didžiojo universitetas, 2022

    Two news portals were selected for comparable corpora building: the Lithuanian portal DELFI and the English portal The Guardian. The compiled corpora comprise 135 Lithuanian articles from DELFI portal and 135 English articles from the Guardian portal. The main criterion for article extraction from the portals was the presence of the two keywords in the articles: "vaccination" and "vaccine". The selected time period for the articles was from January 2021 to September 2021. 30 (15 Lithuanian and 15 English) articles were selected for each month of this period. The extracted articles were used to build two types of comparable corpora necessary for further analysis: full–text corpora composed of full texts of articles and extract corpora composed of titles and lead paragraphs of articles. The sizes of the full-text corpora are: the Lithuanian full-text corpus ‘Lithuanian media articles on vaccination’ contains 45,827 words, the English full-text corpus ‘English media articles on vaccination’ contains 96,759 words. The sizes of the extract corpora are: the Lithuanian extract corpus contains 4,863 words, the English extract corpus contains 3,828 words.

      32
  • dataset[2022][D][H004]
    Utka, Andrius
    ;
    ; ; ; ;
    Vytauto Didžiojo universitetas, 2022

    English-Lithuanian parallel corpus DVITAS includes original English texts on cybersecurity and their Lithuanian translations aligned on the sentence level. The corpus was compiled for the bilingual terminology extraction project together with English-Lithuanian comparable corpus. The parallel corpus includes the EU legal acts and other documents from the time period of 2006-2021. The documents have been extracted from the EUR-Lex database and other EU institutional repositories. There are 80 aligned files in TMX format in English and Lithuanian, as well as 160 raw files (80 in English, and 80 in Lithunian) in the dataset. The total size of the corpus is 1.4m words (EN-0.77m; LT-0.63m). The corpus contains 35,415 aligned segments.

      93
  • English-Lithuanian parallel corpus DVITAS v2 includes original English texts on cybersecurity and their Lithuanian translations aligned on the sentence level. Version 1 of the corpus was compiled for the bilingual terminology extraction project DVITAS together with English-Lithuanian comparable corpus. The current 2nd version of the corpus features expansion of the 1st version containing additional 27 files and metadata information. The parallel corpus includes the EU legal acts and other documents from the time period of 2006-2022. The documents have been extracted from the EUR-Lex database and other EU institutional repositories. There are 107 aligned files in TMX format in English and Lithuanian, as well as 214 raw files (107 in English, and 107 in Lithuanian) within the dataset. The total size of the corpus is 1.97m words (EN-1.08m; LT-0.88m). The corpus contains 53,792 aligned segments.

      14
  • dataset[2025][H004];
    CLARIN-LT digital library in the Republic of Lithuania, 2025-10-15

    English-Lithuanian Parallel Migration Corpus includes original English texts and their Lithuanian translations, aligned at the sentence level. The texts are drawn from EU legal acts and other migration-related documents published in the EUR-Lex database between 1998 and 2024. The total size of the corpus is 1,223,350 words (EN - 688,410 words; LT - 534,940 words). The corpus contains 43,345 aligned segments (sentences). Within the dataset, the following files are included: 1) EN-LT_Parallel_Migration_Corpus_TMX.zip This file is composed of 51 files in TMX (translation memory exchange) format: - 50 separate EN-LT TMX files with aligned texts - 1 combined file consolidating all 50 EN-LT TMX files 2) EN-LT_Parallel_Migration_Corpus_VERT.zip This file is composed of 102 files in VERT (vertical text) format: - 50 separate EN files with morphological annotation - 1 combined EN file consolidating all 50 EN VERT files - 50 separate LT files with morphological annotation - 1 combined LT file consolidating all 50 LT VERT files Sentence aglinment: Each block corresponds to a TMX translation unit . Morphological annotation structure: EN: wordform | tag | lempos (EN TreeTagger) LT: wordform | lempos | tag (LT MULTEXT-East) Tagset references: https://www.sketchengine.eu/english-treetagger-pipeline-2/ https://www.sketchengine.eu/lithuanian-multext-east-part-of-speech-tagset/ 3) EN-LT_Parallel_Migration_Corpus_TXT.zip This files is composed of 100 files in TXT (plain text) format: - 50 separate EN files - 50 separate LT files 4) EN-LT_Parallel_Migration_Corpus_CSV(Metadata).zip This file is composed of 2 files with metadata in CSV (comma separated values) format: - 1 EN file with metadata - 1 LT file with metadata Metadata categories: Form of document, File name (CELEX number of document), Title of document, Author of document (Institution), Year of Publication, Word count, URL. The dataset comprises a total of 255 files, all ecoded in UTF-8.

      29
  • dataset[2018][D][H004];
    Vytauto Didžiojo universitetas, 2018

    274,460 word corpus comprised of selected primary and secondary law acts of the EU of the period 2015-2017. The corpus was compiled of documents containing words with the root "teis-" (en. law). All of the included documents were extracted from EUR-Lex database.

      21
  • Item type:Dataset,
    Lithuanian-English Cybersecurity Termbase v.0.1
    [Lietuvių-anglų kalbų kibernetinio saugumo terminų bazė v.0.1]
    dataset[2023][D][H004]
    Utka, Andrius
    ;
    ; ; ; ;
    Vytauto Didžiojo universitetas, 2023

    The bilingual termbase is TBX export of the online termbase https://www.terminologue.org/csterms/. The termbase includes terms for 233 cybersecurity concepts.

      47
  • dataset[2024][D][H004]
    Armaselu, Florentina
    ;
    Liebeskind, Chaya
    ;
    Marongiu, Paola
    ;
    McGillivray, Barbara
    ;
    ;
    Apostol, Elena Simona
    ;
    Truică, Ciprian-Octavian
    European Organization for Nuclear Research, 2024

    LLODIA (Linguistic Linked Open Data for Diachronic Analysis) model developed within the Nexus Linguarum WG4 UC4.2.1 use case in humanities.

      24  5
  • dataset[2024][D][H004];
    Utka, Andrius
    Mykolo Romerio universitetas, 2024-10-04

    The data is provided in two files: one containing questionnaire-data and the other containing the respondentents' data. The questionnaire data is in a TXT file, which includes the survey questions and possible responses. The respondents' data is in a TSV file with 26 columns, detailing anonymised respondent information and their answers: Timestamp, Agreement, Region, Age, Education, Profession, Preference of the 1st term, Reasons, ..., Preference of 10th term, Reasons. A total of 593 respondents participated, representing diverse age groups, regions, and levels of expertise. Participants were asked to choose the most appropriate Lithuanian terms for 10 cybersecurity concepts (cyberattack, spam, denial-of-service attack, man-in-the-middle attack, brute force attack, phishing, botnet, hacker, honeypot method, zero-day vulnerability). They could either select term provided in the questionnaire or suggest their own, giving reasons for their selections. The dataset facilitates research into terminology preferences, revealing which types of terms are preferred by users (borrowings, metaphorical calques, or descriptive terms) and how preferences vary across two respondent groups: students versus graduates and cybersecurity experts versus the general public. Additionally, data on respondents' reasoning revealed key factors in determining term suitability. The self-suggested terms underscore respondents' creative potential and their strong interest in maintaining national terminology. Rackevičienė, Sigita & Utka, Andrius (2024) "Preferences of Lithuanian cybersecurity synonymous terms in different user groups." Kalbų studijos 44, 107-122.

      16