Skip to content

CV

Employment

2026–present
Assistant Professor, Aarhus University
Research on continuous development and evaluation of language models · Manage AarhusNLP
2025–2026
Postdoc, Aarhus University
Research on continuous development and evaluation of language models

Education

2021–2024
PhD, Aarhus University
Dissertation: Evaluating and Learning Representations for Language and Genetics
Center for Humanities Computing, in collaboration with Quantitative Genomics Group and Aarhus University Hospital
Main supervisor: Kristoffer Nielbo · Co-supervisors: Doug Speed, Andreas Danielsen
Research stays: UCLA (2023, Prof. Vwani Roychowdhury), UC Berkeley (2023, Prof. Tim Tangherlini)
2016–2022
BSc & MSc Cognitive Science, Aarhus University
Elective: Mathematics · GPA: 11.67/12.00

Professional Experience

2024–2025
Research Assistant, Aarhus University
Teaching and research in Natural Language Processing at Cognitive Science
2018–2022
Instructor, Aarhus University
Natural Language Processing, Computational Modelling, and Experimental Methods at Cognitive Science
Topics: GLM, GLMM, Bayesian modelling, R, Python, HPC, NLP, cognitive modelling
2018–2021
Student Developer, Center for Humanities Computing Aarhus
HPC, NLP, and information extraction
2017–2020
Junior Consultant, JHN Processor
Data management, data collection, economics, and user experience

Funding

2022–2025
Multiple Grants, Danish E-Infrastructure Cooperation
>300,000 GPU core hours and >1,000,000 CPU core hours
Case numbers: DeiC-AU-N5-2024079, DeiC-AU-N1-2025144, DeiC-KU-N5-2025117, H2-2023-15, H2-2023-16, 2022-H2-11

Co-wrote the following grant applications (not grant holder)

2025–2028
Lex.llm, The Augustinus Foundation and the Aage and Johanne Louis-Hansen Foundation
2025–2027
EuropeanCity2, Horizon Europe (EU)
2022–present
Danish Foundation Models, Ministry of Digital Affairs and Danish e-Infrastructure Cooperation

Peer Review

Active reviewer
ACL Rolling Review (ARR), NeurIPS, ICLR, and the Journal of Open Source Software (JOSS)
Previously reviewed for
Northern European Journal of Language Technology, Acta Psychiatrica Scandinavica, Behavior Research Methods, Computational Humanities Research, and Digital Humanities Benelux

Counseling

2023-2026
The Danish Agency for Digital Governance (Digitaliseringsstyrelsen)
Invited presentations and counselling on data and evaluation of AI
2026
Agency for Climate Data (Klimadatastyrelsen)
Invited presentations measuring the impact of AI

Supervision

2025–present
Jakob Grøhn Damgaard
PhD Co-supervisor
2025
Anton Drasbæk
Master's thesis supervisor
2025
Jørgen Højlund Wibe
Master's thesis supervisor
2023
Emil Jessen
Master's thesis supervisor

Talks

Selected talks and presentations

2026
AI-Arenaen: Hvordan klarer sprogmodellerne sig på dansk?
Driving AI 2026, IDA (The Danish Society of Engineers)
2024
Workshop: Danish NLP
Danish Digitalization, Data Science and AI 1.0 (D3A), Nyborg, Denmark
2023
AI4Welfare: Dansk sprogteknologi til bedre velfærd
Kommunernes Landsforening (KL), Denmark
2023
Danish Foundation Models: Validerede sprogmodeller til dansk
Sprogteknologisk Konference 2023
2023
An Intuitive Introduction to Developments in Machine Learning and Language Technology
KMD Insight '23
2022
Danish Foundation Models and Danish Open Source Models
Danish Data Science 2022, Billund, Denmark
2022
2021

Non-scientific Publications

2026
Danmarks AI-visioner kræver data, vi ikke har: Her er tre veje til at skaffe dem, Version2
Op-ed on the data foundation required for Danish AI
2026
Open source er en strategisk nødvendighed i AI: Sådan kommer vi med i kapløbet, DataTech
Co-authored op-ed on the importance of open source for Danish AI independence
2025
AI's klimaaftryk: Hvad ved vi, og hvad gætter vi os til, Version2
Co-authored op-ed on the climate footprint of AI

Open-source Projects

Selected open source projects

2025-present
Dynaword
A framework for continuously updated, openly licensed text corpora, with editions for Danish, Norwegian, Swedish, Icelandic, Faroese, and Dutch · All subsets are (2026) the largest corpus of openly licensed text for their respective language · Core developer and maintainer
2024-present
Massive Multilingual Embedding Benchmark (MTEB)
The de-facto Python package and benchmark for evaluating text, image, audio and video embedding models across languages and use cases · Core developer and maintainer
2024-2025
Scandinavian Embedding Benchmark
The de-facto Benchmark for estimating the quality of Scandinavian embedding model. Later merged into MTEB · Core developer and maintainer
2023
DFM Sentence Encoder (large)
A Danish sentence embedding model, still (2025) Pareto optimal on EuroEval (Danish, NLU)
2023-present
Augmenty
A Python package for text augmentation with use cases in bias detection, evaluating model robustness, and improving model performance · Core developer and maintainer
2023-present
timeseriesflattener
A package for converting irregularly spaced time series, such as electronic health records, into statically shaped data frames · Initial developer, maintained by others
2022-present
TextDescriptives
A package for extracting text features such as dependency dynamics and metrics of text quality · Co-developer and maintainer
2022-present
Tomsup 👍
Theory of Mind Simulation using Python · Agent-based simulation implementing variational recursive k-ToM · Core developer and maintainer
2022-present
UD_Danish-DDT
The Danish Universal Dependencies Treebank, a high quality linguistic resource · Maintainer
2021-present
DaCy
State-of-the-art Danish NLP · POS tagging (98.37 acc), NER (84.39 F1), dependency parsing (88.44 LAS) on DDT and DaNE · Core developer and maintainer
2021-present
DANSK
DANSK: Danish Annotations for NLP Specific TasKs is a dataset consisting of texts from diverse domains annotated for 18 entities. Actively used in EuroEval · Language resource

Open-source Contributions

Selected contributions

2026
Danish ASR Leaderboard
Added uncertainty estimates and improved evaluation metrics
2026
text2num, Allo-Media
Added support for Danish
2025-2026
ComparIA, French Government (beta.gouv.fr)
Danish translations and LaTeX rendering support for the open LLM arena
2025
spacy-lookup-data, Explosion
Added Danish Lexeme probabilities
2024
datasets, Huggingface
Fixes for compatibility issue with numpy >=2.0.0
2024
curated-transformers, Explosion
Added support for ELECTRA models
2024
spacy-curated-transformers, Explosion
Added support for ELECTRA tokenizers
2023
confection, Explosion
Fixed issue where config where could not be filled
2023
curated-transformers, Explosion
Added support for ELECTRA models
2022
transformers, Huggingface
Bugfixes for training masked language models using flax
2021
spacy-transformers, Explosion
Allow passing arguments to the transformer backend to obtain attention weights