FiHuCoMet corpus

FiHuComet Corpus

2026.02.28.
FiHuComet Corpus

The FiHuComet corpus is a Finnish and Hungarian metaphor corpus in which potentially metaphorical structures were annotated using the MetaID procedure. We applied the method not only to Hungarian, but also adapted it to Finnish, making the corpus bilingual. However, it cannot be considered a parallel corpus in the classical sense, as the texts it contains are not translations of each other, but Hungarian and Finnish news texts reporting on the same topics and events, and can therefore be defined as a "quasi" parallel corpus.

The texts selected for the corpus were short news articles. The selection of this text type was based on the assumption that these shorter news texts focus primarily on presenting events and contain fewer linguistic expressions of individual attitudes, thus making them more objective and fact-oriented. The texts were selected from reports on events of international significance, as these were covered by both the Hungarian and Finnish press, which made it easier to find texts with the same subject matter in both languages.

The criteria for selection into the corpus were that, in addition to thematic relevance, the news items had to date from the same period. The texts come from two independent news portals that are freely accessible online. Sampling and annotation took place in two periods: the corpus was created and the first texts were processed in February 2022, while further expansion took place in October 2023.

The Hungarian subcorpus contains 4829 words, while the Finnish subcorpus contains 5824 words. The corpus is considered small in size, but it is manually annotated, meaning that every potentially metaphorical structure has been analyzed by a human annotator. The purpose of creating the corpus was, on the one hand, to conduct a comparative, pilot study of the metaphorical patterns of the two languages and, on the other hand, to test the adaptation of the MetaID method to the Finnish language.

Based on the results of the small-scale corpus analysis, it appeared that in Hungarian texts, metaphorical structures are concentrated around verb forms, and the secondary figures (landmarks) of verbs considered metaphorical and the circumstances of the represented event were given more detailed lexical elaborations, which contributed to metaphorization. In contrast, in the Finnish texts, the primary participants contributed more significantly to the elaboration of the metaphors. Overall, in terms of relative frequency, potentially metaphorical expressions occurred at a higher rate in the Finnish texts. In the processed corpus of both languages, we often found metaphorization under the word, and we could not detect any significant difference between the two subcorpora; both Finnish and Hungarian metaphorical structures were characterized by the use of postpositions and suffixes (the relative frequency of the latter was prominent) for elaboration.

 

Publications:

 

Further information: bajzat.timea@btk.elte.hu