CV¶
Employment¶
2026–present
Assistant Professor, Aarhus University
Research on continuous development and evaluation of language models · Manage AarhusNLP
Research on continuous development and evaluation of language models · Manage AarhusNLP
2025–2026
Postdoc, Aarhus University
Research on continuous development and evaluation of language models
Research on continuous development and evaluation of language models
Education¶
2021–2024
PhD, Aarhus University
Dissertation: Evaluating and Learning Representations for Language and Genetics
Center for Humanities Computing, in collaboration with Quantitative Genomics Group and Aarhus University Hospital
Main supervisor: Kristoffer Nielbo · Co-supervisors: Doug Speed, Andreas Danielsen
Research stays: UCLA (2023, Prof. Vwani Roychowdhury), UC Berkeley (2023, Prof. Tim Tangherlini)
Dissertation: Evaluating and Learning Representations for Language and Genetics
Center for Humanities Computing, in collaboration with Quantitative Genomics Group and Aarhus University Hospital
Main supervisor: Kristoffer Nielbo · Co-supervisors: Doug Speed, Andreas Danielsen
Research stays: UCLA (2023, Prof. Vwani Roychowdhury), UC Berkeley (2023, Prof. Tim Tangherlini)
2016–2022
BSc & MSc Cognitive Science, Aarhus University
Elective: Mathematics · GPA: 11.67/12.00
Elective: Mathematics · GPA: 11.67/12.00
Professional Experience¶
2024–2025
Research Assistant, Aarhus University
Teaching and research in Natural Language Processing at Cognitive Science
Teaching and research in Natural Language Processing at Cognitive Science
2018–2022
Instructor, Aarhus University
Natural Language Processing, Computational Modelling, and Experimental Methods at Cognitive Science
Topics: GLM, GLMM, Bayesian modelling, R, Python, HPC, NLP, cognitive modelling
Natural Language Processing, Computational Modelling, and Experimental Methods at Cognitive Science
Topics: GLM, GLMM, Bayesian modelling, R, Python, HPC, NLP, cognitive modelling
2018–2021
Student Developer, Center for Humanities Computing Aarhus
HPC, NLP, and information extraction
HPC, NLP, and information extraction
2017–2020
Junior Consultant, JHN Processor
Data management, data collection, economics, and user experience
Data management, data collection, economics, and user experience
Funding¶
2022–2025
Multiple Grants, Danish E-Infrastructure Cooperation
>300,000 GPU core hours and >1,000,000 CPU core hours
Case numbers: DeiC-AU-N5-2024079, DeiC-AU-N1-2025144, DeiC-KU-N5-2025117, H2-2023-15, H2-2023-16, 2022-H2-11
>300,000 GPU core hours and >1,000,000 CPU core hours
Case numbers: DeiC-AU-N5-2024079, DeiC-AU-N1-2025144, DeiC-KU-N5-2025117, H2-2023-15, H2-2023-16, 2022-H2-11
Co-wrote the following grant applications (not grant holder)
2025–2028
Lex.llm, The Augustinus Foundation and the Aage and Johanne Louis-Hansen Foundation
2025–2027
EuropeanCity2, Horizon Europe (EU)
2022–present
Danish Foundation Models, Ministry of Digital Affairs and Danish e-Infrastructure Cooperation
Peer Review¶
Active reviewer
ACL Rolling Review (ARR), NeurIPS, ICLR, and the Journal of Open Source Software (JOSS)
ACL Rolling Review (ARR), NeurIPS, ICLR, and the Journal of Open Source Software (JOSS)
Previously reviewed for
Northern European Journal of Language Technology, Acta Psychiatrica Scandinavica, Behavior Research Methods, Computational Humanities Research, and Digital Humanities Benelux
Northern European Journal of Language Technology, Acta Psychiatrica Scandinavica, Behavior Research Methods, Computational Humanities Research, and Digital Humanities Benelux
Counseling¶
2023-2026
The Danish Agency for Digital Governance (Digitaliseringsstyrelsen)
Invited presentations and counselling on data and evaluation of AI
Invited presentations and counselling on data and evaluation of AI
2026
Agency for Climate Data (Klimadatastyrelsen)
Invited presentations measuring the impact of AI
Invited presentations measuring the impact of AI
Supervision¶
2025–present
Jakob Grøhn Damgaard
PhD Co-supervisor
PhD Co-supervisor
2025
Anton Drasbæk
Master's thesis supervisor
Master's thesis supervisor
2025
Jørgen Højlund Wibe
Master's thesis supervisor
Master's thesis supervisor
2023
Emil Jessen
Master's thesis supervisor
Master's thesis supervisor
Talks¶
Selected talks and presentations
2026
AI-Arenaen: Hvordan klarer sprogmodellerne sig på dansk?
Driving AI 2026, IDA (The Danish Society of Engineers)
Driving AI 2026, IDA (The Danish Society of Engineers)
2024
2023
AI4Welfare: Dansk sprogteknologi til bedre velfærd
Kommunernes Landsforening (KL), Denmark
Kommunernes Landsforening (KL), Denmark
2023
Danish Foundation Models: Validerede sprogmodeller til dansk
Sprogteknologisk Konference 2023
Sprogteknologisk Konference 2023
2023
An Intuitive Introduction to Developments in Machine Learning and Language Technology
KMD Insight '23
KMD Insight '23
2022
2022
When Norwegians are Better than Danes at Danish: The State and Shortcomings of Danish NLP, and How We Fix Them
AU Digital Innovation Conference, Aarhus, Denmark
AU Digital Innovation Conference, Aarhus, Denmark
2021
Non-scientific Publications¶
2026
Danmarks AI-visioner kræver data, vi ikke har: Her er tre veje til at skaffe dem, Version2
Op-ed on the data foundation required for Danish AI
Op-ed on the data foundation required for Danish AI
2026
Open source er en strategisk nødvendighed i AI: Sådan kommer vi med i kapløbet, DataTech
Co-authored op-ed on the importance of open source for Danish AI independence
Co-authored op-ed on the importance of open source for Danish AI independence
2025
AI's klimaaftryk: Hvad ved vi, og hvad gætter vi os til, Version2
Co-authored op-ed on the climate footprint of AI
Co-authored op-ed on the climate footprint of AI
Open-source Projects¶
Selected open source projects
2025-present
Dynaword
A framework for continuously updated, openly licensed text corpora, with editions for Danish, Norwegian, Swedish, Icelandic, Faroese, and Dutch · All subsets are (2026) the largest corpus of openly licensed text for their respective language · Core developer and maintainer
A framework for continuously updated, openly licensed text corpora, with editions for Danish, Norwegian, Swedish, Icelandic, Faroese, and Dutch · All subsets are (2026) the largest corpus of openly licensed text for their respective language · Core developer and maintainer
2024-present
Massive Multilingual Embedding Benchmark (MTEB)
The de-facto Python package and benchmark for evaluating text, image, audio and video embedding models across languages and use cases · Core developer and maintainer
The de-facto Python package and benchmark for evaluating text, image, audio and video embedding models across languages and use cases · Core developer and maintainer
2024-2025
Scandinavian Embedding Benchmark
The de-facto Benchmark for estimating the quality of Scandinavian embedding model. Later merged into MTEB · Core developer and maintainer
The de-facto Benchmark for estimating the quality of Scandinavian embedding model. Later merged into MTEB · Core developer and maintainer
2023
DFM Sentence Encoder (large)
A Danish sentence embedding model, still (2025) Pareto optimal on EuroEval (Danish, NLU)
A Danish sentence embedding model, still (2025) Pareto optimal on EuroEval (Danish, NLU)
2023-present
Augmenty
A Python package for text augmentation with use cases in bias detection, evaluating model robustness, and improving model performance · Core developer and maintainer
A Python package for text augmentation with use cases in bias detection, evaluating model robustness, and improving model performance · Core developer and maintainer
2023-present
timeseriesflattener
A package for converting irregularly spaced time series, such as electronic health records, into statically shaped data frames · Initial developer, maintained by others
A package for converting irregularly spaced time series, such as electronic health records, into statically shaped data frames · Initial developer, maintained by others
2022-present
TextDescriptives
A package for extracting text features such as dependency dynamics and metrics of text quality · Co-developer and maintainer
A package for extracting text features such as dependency dynamics and metrics of text quality · Co-developer and maintainer
2022-present
Tomsup 👍
Theory of Mind Simulation using Python · Agent-based simulation implementing variational recursive k-ToM · Core developer and maintainer
Theory of Mind Simulation using Python · Agent-based simulation implementing variational recursive k-ToM · Core developer and maintainer
2022-present
UD_Danish-DDT
The Danish Universal Dependencies Treebank, a high quality linguistic resource · Maintainer
The Danish Universal Dependencies Treebank, a high quality linguistic resource · Maintainer
2021-present
DaCy
State-of-the-art Danish NLP · POS tagging (98.37 acc), NER (84.39 F1), dependency parsing (88.44 LAS) on DDT and DaNE · Core developer and maintainer
State-of-the-art Danish NLP · POS tagging (98.37 acc), NER (84.39 F1), dependency parsing (88.44 LAS) on DDT and DaNE · Core developer and maintainer
2021-present
DANSK
DANSK: Danish Annotations for NLP Specific TasKs is a dataset consisting of texts from diverse domains annotated for 18 entities. Actively used in EuroEval · Language resource
DANSK: Danish Annotations for NLP Specific TasKs is a dataset consisting of texts from diverse domains annotated for 18 entities. Actively used in EuroEval · Language resource
Open-source Contributions¶
Selected contributions
2026
2026
2025-2026
ComparIA, French Government (beta.gouv.fr)
Danish translations and LaTeX rendering support for the open LLM arena
Danish translations and LaTeX rendering support for the open LLM arena
2025
spacy-lookup-data, Explosion
Added Danish Lexeme probabilities
Added Danish Lexeme probabilities
2024
datasets, Huggingface
Fixes for compatibility issue with numpy >=2.0.0
Fixes for compatibility issue with numpy >=2.0.0
2024
curated-transformers, Explosion
Added support for ELECTRA models
Added support for ELECTRA models
2024
spacy-curated-transformers, Explosion
Added support for ELECTRA tokenizers
Added support for ELECTRA tokenizers
2023
confection, Explosion
Fixed issue where config where could not be filled
Fixed issue where config where could not be filled
2023
curated-transformers, Explosion
Added support for ELECTRA models
Added support for ELECTRA models
2022
transformers, Huggingface
Bugfixes for training masked language models using flax
Bugfixes for training masked language models using flax
2021
spacy-transformers, Explosion
Allow passing arguments to the transformer backend to obtain attention weights
Allow passing arguments to the transformer backend to obtain attention weights