Meta tags:
description= :book: A curated list of resources dedicated to Natural Language Processing (NLP) - keon/awesome-nlp;
Headings (most frequently used words):
and, nlp, in, models, libraries, embeddings, datasets, language, data, uh, oh, navigation, text, saved, searches, awesome, files, research, tutorials, multilingual, license, footer, reading, information, extraction, languages, corpora, search, code, repositories, users, issues, pull, requests, provide, feedback, keon, menu, use, to, filter, your, results, more, quickly, folders, latest, commit, history, repository, scope, contents, summaries, trends, prominent, labs, tasks, methods, frameworks, for, per, see, also, citation, about, releases, packages, contributors, historical, highlights, content, videos, online, courses, books, services, annotation, tools, tokenization, morphology, segmentation, pos, tagging, dependency, parsing, named, entity, recognition, coreference, resolution, classification, sentiment, analysis, topic, modeling, summarization, machine, translation, question, answering, comprehension, beyond, ner, retrieval, speech, pretraining, adaptation, cross, lingual, evaluation, benchmarks, reasoning, test, time, compute, long, context, alternative, architectures, factuality, hallucination, calibration, probing, interpretability, efficient, small, instruction, tuning, preference, optimization, bias, fairness, safety, arabic, chinese, anthology, danish, dutch, german, hungarian, indic, treebanks, word, tooling, indonesian, korean, blogs, persian, polish, portuguese, spanish, thai, ukrainian, urdu, uzbek, vietnamese, other, topics, resources, contributing, stars, watchers, forks, that, need, login, access, can, be, gained, via, email,
Text of the page (most frequently used words):
and (360), for (250), nlp (233), language (124), with (108), the (103), #models (100), open (78), languages (71), natural (65), 2024 (64), text (63), top (63), processing (61), 2025 (61), back (54), based (47), library (44), multilingual (43), model (42), data (40), python (37), from (36), embeddings (35), across (33), ner (33), tasks (33), datasets (32), word (32), learning (30), llms (29), reasoning (28), retrieval (27), toolkit (27), pos (26), annotation (26), korean (25), evaluation (25), benchmark (25), level (25), corpus (24), free (24), long (24), libraries (23), tagging (23), persian (23), bert (22), context (22), neural (22), llm (21), speech (21), machine (20), chinese (20), instruction (20), tokenization (20), lms (20), extraction (20), research (20), arabic (20), resources (19), deep (19), parsing (19), classification (19), analysis (19), code (19), generation (18), large (18), vietnamese (18), dependency (18), translation (18), multi (18), training (18), topic (18), github (17), awesome (17), tools (17), sentiment (17), hindi (17), pretraining (17), english (16), treebank (16), 2023 (16), spanish (16), encoder (16), source (16), 2026 (15), list (15), summarization (15), transformer (15), modeling (15), that (14), time (14), corpora (14), covering (14), sentence (14), strong (14), trained (14), embedding (14), coreference (14), document (14), tool (14), curated (13), rag (13), see (13), dataset (13), package (13), efficient (13), platform (13), small (13), tuning (13), transformers (13), attention (13), modern (13), information (12), also (12), linguistic (12), indonesian (12), tokens (12), scale (12), fast (12), recognition (12), entity (12), cross (12), fine (12), search (12), token (12), test (12), all (12), your (12), end (12), this (11), including (11), task (11), benchmarks (11), indic (11), quality (11), parser (11), named (11), inference (11), built (11), you (10), can (10), security (10), production (10), collection (10), thai (10), implementation (10), other (10), family (10), pretrained (10), segmentation (10), foundation (10), word2vec (10), course (10), high (10), spacy (10), low (10), state (10), features (10), view (10), foundational (10), such (10), systems (9), tokenizer (9), bilingual (9), sentences (9), universal (9), tuned (9), java (9), methods (9), morphological (9), pytorch (9), dutch (9), support (9), over (9), resource (9), framework (9), bias (9), more (9), art (9), self (9), decoder (9), sparse (9), hallucination (9), like (9), labeling (9), license (8), general (8), current (8), using (8), standard (8), first (8), format (8), llama (8), university (8), via (8), common (8), human (8), shot (8), mteb (8), google (8), web (8), are (8), performance (8), architecture (8), 200 (8), lingual (8), api (8), error (7), please (7), reload (7), asian (7), words (7), semantic (7), sequence (7), urdu (7), portuguese (7), used (7), analyzer (7), scala (7), nltk (7), preference (7), feedback (7), huggingface (7), aware (7), dense (7), unified (7), survey (7), factuality (7), calibration (7), detection (7), question (7), answering (7), compute (7), stanford (7), community (6), contributing (6), topics (6), prompt (6), generative (6), here (6), 100 (6), news (6), southeast (6), relation (6), available (6), part (6), fully (6), dependencies (6), suite (6), samples (6), pipeline (6), danish (6), moe (6), mmlu (6), domain (6), alignment (6), safety (6), matches (6), widely (6), written (6), without (6), interpretable (6), representations (6), probing (6), gpt (6), linear (6), sota (6), system (6), following (6), nlu (6), seq2seq (6), baseline (6), reading (6), how (6), tokenizers (6), building (6), teams (6), various (6), clojure (6), node (6), enterprise (6), navigation (5), there (5), was (5), while (5), repository (5), keon (5), serving (5), prompting (5), augmented (5), papers (5), classical (5), japanese (5), parallel (5), core (5), 400 (5), fasttext (5), covers (5), roberta (5), only (5), wikipedia (5), wiki (5), advanced (5), blog (5), tutorials (5), browser (5), alternative (5), scratch (5), elmo (5), iit (5), german (5), lemmatization (5), deepseek (5), book (5), per (5), many (5), resolution (5), sft (5), reference (5), optimization (5), parameter (5), structured (5), few (5), settings (5), memory (5), larger (5), distillation (5), autoencoders (5), interpretability (5), scaling (5), dynamic (5), pipelines (5), hybrid (5), meta (5), beyond (5), simple (5), loaders (5), statistical (5), traditional (5), ready (5), bpe (5), subword (5), folia (5), spark (5), projects (5), use (5), online (5), actions (5), content (5), include (5), network (5), creating (5), kotlin (5), group (5), notebooks (5), issues (5), summaries (5), illustrated (5), loading (4), page (4), readme (4), mining (4), dedicated (4), find (4), tooling (4), purpose (4), lists (4), out (4), emnlp (4), tts (4), ukrainian (4), best (4), polish (4), annotated (4), most (4), tags (4), nmt (4), technology (4), style (4), provides (4), polarity (4), facebook (4), review (4), stopwords (4), hungarian (4), focus (4), qwen (4), correction (4), morphology (4), anthropic (4), fairness (4), results (4), zero (4), weight (4), adaptation (4), layer (4), cache (4), quantization (4), bge (4), paragraph (4), accuracy (4), non (4), gemma (4), reproducible (4), than (4), clustering (4), architectures (4), feature (4), transcoders (4), outputs (4), hop (4), knowledge (4), factual (4), form (4), set (4), grained (4), extension (4), matching (4), process (4), step (4), surpasses (4), into (4), questions (4), refresh (4), translate (4), behind (4), work (4), size (4), new (4), functionality (4), files (4), grade (4), asr (4), includes (4), scoring (4), series (4), then (4), pre (4), comprehension (4), networks (4), vectors (4), canonical (4), approaches (4), byte (4), rust (4), wide (4), supports (4), costs (4), powered (4), interface (4), services (4), julia (4), ruby (4), regular (4), metrics (4), visual (4), notable (4), acl (4), pull (4), action (3), terms (3), report (3), stars (3), cc0 (3), about (3), 2018 (3), useful (3), citation (3), techniques (3), template (3), engineering (3), hebrew (3), texts (3), elasticsearch (3), russian (3), good (3), tagger (3), books (3), 250 (3), big (3), constituency (3), sailor (3), mistral (3), chat (3), mit (3), uzbek (3), focused (3), etc (3), multiple (3), center (3), both (3), annotations (3), tagged (3), stemmer (3), squad (3), archive (3), movie (3), institute (3), bahasa (3), collections (3), indobert (3), telugu (3), patna (3), access (3), labelled (3), sense (3), twitter (3), domains (3), aspect (3), bangla (3), classes (3), treebanks (3), developed (3), industrial (3), binding (3), frog (3), alibaba (3), trainable (3), javascript (3), faking (3), policy (3), value (3), reveals (3), jailbreak (3), integrating (3), turn (3), strategies (3), pairs (3), grpo (3), way (3), ai2 (3), post (3), recipe (3), tülu (3), adapting (3), class (3), loss (3), qwen3 (3), 128k (3), competitive (3), much (3), encoders (3), specific (3), sae (3), shows (3), introduces (3), graphs (3), tracing (3), methodology (3), recall (3), learns (3), claim (3), uncertainty (3), before (3), through (3), faithfulness (3), under (3), comprehensive (3), rope (3), extends (3), recurrent (3), tradeoffs (3), patterns (3), minimal (3), mamba (3), reward (3), produce (3), chain (3), thought (3), adaptive (3), math (3), cot (3), short (3), supervised (3), needle (3), haystack (3), but (3), together (3), pro (3), gap (3), between (3), gemini (3), seamlessm4t (3), 101 (3), transfer (3), default (3), optimized (3), deberta (3), still (3), application (3), endangered (3), disambiguation (3), frameworks (3), linguistically (3), consistent (3), fineweb (3), gensim (3), fields (3), late (3), interaction (3), colbert (3), improves (3), supporting (3), dpr (3), analyzing (3), read (3), real (3), practical (3), bleu (3), chrf (3), learned (3), moses (3), contextual (3), latent (3), box (3), classifier (3), higher (3), cost (3), types (3), flair (3), stanza (3), vocabulary (3), input (3), unicode (3), implementations (3), unigram (3), agnostic (3), well (3), glove (3), where (3), environment (3), xml (3), shoonya (3), cloud (3), intent (3), syntax (3), working (3), variety (3), engine (3), archived (3), snips (3), software (3), opennlp (3), exploring (3), developing (3), applications (3), expression (3), forms (3), jupyter (3), lectures (3), recent (3), labs (3), penn (3), trends (3), repositories (3), commit (3), name (3), insights (3), requests (3), signed (3), another (3), tab (3), window (3), session (3), sign (3), saved (3), documentation (3), copilot (3), solutions (3), explore (3), perform (2), not (2), manage (2), status (2), privacy (2), footer (2), releases (2), forks (2), adjacent (2), scope (2), ancient (2), 10k (2), device (2), vieneu (2), grounded (2), speeches (2), prime (2), minister (2), tensorflow (2), character (2), coverage (2), national (2), reliable (2), uppsala (2), translated (2), 250k (2), suitable (2), french (2), petition (2), chatbot (2), naver (2), blogs (2), klue (2), native (2), mecab (2), multimodal (2), hub (2), providing (2), standardized (2), wordnet (2), sea (2), lion (2), cendol (2), indolem (2), indonlu (2), algorithm (2), sarvam (2), sanskrit (2), bengali (2), whole (2), ulmfit (2), marathi (2), need (2), email (2), labels (2), absa (2), reviews (2), smaller (2), strength (2), kaldi (2), anthology (2), jieba (2), released (2), alongside (2), arabert (2), organized (2), cases (2), when (2), vector (2), mitigation (2), refusal (2), attacks (2), failures (2), user (2), measurement (2), gender (2), masked (2), instruct (2), rlhf (2), less (2), adopted (2), direct (2), instructions (2), follow (2), flan (2), lora (2), peft (2), rank (2), quantized (2), throughput (2), icml (2), precision (2), relevant (2), fastfit (2), prompts (2), setfit (2), apple (2), intelligence (2), thinking (2), technical (2), local (2), global (2), yarn (2), microsoft (2), ones (2), phi (2), distilbert (2), identifies (2), circuits (2), highly (2), inputs (2), skip (2), applies (2), attribution (2), case (2), enables (2), circuit (2), causal (2), computational (2), examples (2), behavior (2), browsing (2), extracting (2), monosemanticity (2), structural (2), what (2), taxonomy (2), choice (2), pattern (2), sampling (2), ssm (2), combining (2), degradation (2), contexts (2), term (2), memorize (2), historical (2), lost (2), middle (2), windows (2), position (2), rotary (2), length (2), aime (2), baselines (2), verification (2), discriminative (2), clipping (2), four (2), improvements (2), mcts (2), systematic (2), study (2), openai (2), pure (2), tree (2), consistency (2), improve (2), direction (2), explicit (2), nested (2), score (2), longbench (2), synthetic (2), exposing (2), diverse (2), evidence (2), critique (2), beir (2), flores (2), xnli (2), xtreme (2), african (2), speaker (2), madlad (2), nllb (2), 500 (2), cohere (2), massively (2), aya (2), weights (2), base (2), substrate (2), denoising (2), bart (2), modernbert (2), nli (2), since (2), around (2), agents (2), kits (2), uralic (2), some (2), uralicnlp (2), offers (2), server (2), conll (2), udpipe (2), generate (2), edu (2), redpajama (2), index (2), pointer (2), mmteb (2), query (2), trace (2), retrieve (2), one (2), gritlm (2), nomic (2), grammatical (2), srl (2), rebel (2), triviaqa (2), era (2), atlas (2), extractive (2), metric (2), ter (2), sacrebleu (2), facto (2), comet (2), fairseq (2), understanding (2), abstractive (2), generator (2), hierarchical (2), regularized (2), bigartm (2), lda (2), dirichlet (2), allocation (2), white (2), pyss3 (2), sst (2), category (2), coref (2), span (2), handles (2), crf (2), parsers (2), reducing (2), complex (2), theory (2), tokenized (2), sentencepiece (2), unsupervised (2), lemma (2), leaderboards (2), contextualized (2), grams (2), explainer (2), agent (2), control (2), label (2), argilla (2), flat (2), users (2), lab (2), management (2), easy (2), hosted (2), active (2), freemium (2), collaborative (2), discourse (2), offering (2), doccano (2), brat (2), rapid (2), custom (2), cognitive (2), service (2), amazon (2), related (2), vscode (2), clojurescript (2), visualizations (2), mallet (2), apache (2), epic (2), alike (2), ucto (2), bllip (2), crfsuite (2), conditional (2), random (2), crfs (2), sequential (2), evaluate (2), dsl (2), allows (2), which (2), regex (2), industry (2), build (2), very (2), extensible (2), modelling (2), jptdp (2), interactive (2), better (2), plain (2), adversarial (2), world (2), beginner (2), friendly (2), hands (2), hugging (2), face (2), authors (2), program (2), rnns (2), taking (2), introduction (2), carnegie (2), mellon (2), oxford (2), videos (2), courses (2), paper (2), guides (2), explanations (2), guide (2), sebastian (2), ruder (2), newsletter (2), playbook (2), engineer (2), their (2), contributions (2), development (2), create (2), project (2), prominent (2), rolling (2), venues (2), they (2), add (2), pull_request_template (2), assets (2), 622 (2), commits (2), last (2), message (2), menu (2), appearance (2), cancel (2), searches (2), business (2), developer (2), customer (2), devops (2), app (2), share, personal, cookies, contact, docs, inc, contributors, packages, published, watching, 568, watchers, activity, note, https, com, url, year, kim, woo, author, title, misc, consider, citing, mlops, modalities, nlph_resources, doing, cltk, lao, icu, pymorphy2, 20m, law, articles, film, subtitles, evb, sql, 2020, findings, vitext2sql, 75m, vntqcorpus, txt, hours, recorded, hcmus, ailab, vivos, ud_vietnamese, bktreebank, vistral, vinai, phogpt, bartpho, phobert, voice, cloning, pyvi, vncorenlp, vitk, underthesea, culturally, rows, soas, urduhack, ukrainianlt, thailand, inter, openthaigpt, scb, 10x, typhoon, wangchanberta, synthai, cutkum, cluster, jtcc, pythainlp, sent2vec, rigochat, bsc, barcelona, supercomputing, salamandra, basque, latxa, bne, beto, compilation, unannotated, billion, copenhagen, columbian, political, detect, censor, clean, profanity, hate, bullying, speaking, countries, spanlp, portulan, albertina, maritaca, sabiá, brazilian, bertimbau, clef, 2008, 2009, hamshahri, syntactically, updt, 982, verbs, valency, lexicon, syntactic, semeval, 2010, perlex, 25m, farsiyar, persianner, 682, iob, armanpersonercorpus, lscp, 120m, 27m, casual, tweets, colloquial, freely, upc, farsi, manually, bijankhan, dorna, persianmind, parsbert, cleaning, virastar, parsianalyzer, partial, perstem, keyphrase, perke, parsivar, hazm, html, korquad, expired, blue, house, site, petitions, major, south, newspaper, chosun, ilbo, korea, science, kaist, kangwon, dsindex, hyperclova, exaone, polyglot, skt, kobert, running, client, side, webassembly, 1mb, offline, garu, kiwi, splitter, kss, konlp, koalanlp, konlpy, seacrowd, dictionary, indosum, idn, 39k, 900k, panl10n, kompas, tempo, ilps, singapore, regional, nusacrowd, indonesia, sastrawi, stemming, pysastrawi, bharatgpt, krutrim, airavata, continuation, openhathi, indictrans2, 2022, indicbert, ai4bharat, indicnlp, fastai, inltk, kannada, sivareddy, python3, port, transliteration, helpers, oscar, albert, bunch, crawl, languge, hindi2vec, tdil, aggregates, lot, otherwise, gated, sentiwordnet, bombay, cfilt, tamil, sail, 2015, login, gained, bbc, 60k, peter, graham, isi, fire, above, mentioned, representational, layered, off, shelf, does, alpino, surface, realiser, simplenlg, simplenlg_nl, danlp, funnlp, baichuan, tsinghua, chatglm3, glm, improved, mlm, macbert, masking, wwm, hit, ltp, hanlp, fudannlp, snownlp, arabicmmlu, aggregated, labr, largest, multidomain, sdaia, allam, jais, araelectra, msa, dialectal, camelbert, qcri, farasa, segmenter, coptic, rftokenizer, pyarabic, jsastem, goarabic, dialect, camel, click, section, expand, occurs, conflicts, internalized, values, steering, reduces, vlaf, conflict, agreement, scripts, indicsafe, modular, defenses, risk, categories, teleai, 4000, dialogues, scenarios, safedialbench, finetuning, narrow, insecure, unexpectedly, produces, broad, unrelated, emergent, misalignment, moderation, wildguard, strategically, complying, during, tailoring, answers, beliefs, sycophancy, toxicity, realtoxicityprompts, demographic, axes, holisticbias, winobias, social, crows, measuring, stereotypical, stereoset, synthesizes, response, aligned, nothing, filtered, subset, official, magpie, dpo, trl, goes, lima, among, simpler, generated, against, constitution, constitutional, 1600, super, naturalinstructions, bootstrapping, instructgpt, finetuned, learners, bundling, prefix, ia3, others, decomposed, dora, adapters, modest, hardware, qlora, tgi, sglang, pagedattention, vllm, portable, gguf, cpp, sensitivity, wise, mixed, improvement, uniform, kv8, kvtuner, activation, awq, gptq, deploying, compact, near, stella, gte, siamese, sharing, bit, qat, reduction, 235b, modes, 30b, a3b, activating, parameters, 27b, ratio, keep, tractable, nope, smollm3, smollm2, distilled, minilm, rather, mechanism, finding, explanation, reconstructing, yield, saes, beat, claude, haiku, rhyme, planning, studies, biology, construct, replacement, interactions, revealing, identifying, driving, influence, functions, neuronpedia, towards, toy, superposition, pyramid, probes, locating, editing, associations, rome, limitations, alternatives, classifiers, belinkov, primer, bertology, trains, reason, generating, gains, biography, factbench, auroc, cure, think, hard, required, responses, rates, persist, even, halluhard, logits, principled, quantification, fact, checking, formally, separates, franq, substantially, worse, calibrated, extended, single, claims, atomic, extrinsic, intrinsic, regeneration, resist, leakage, hallulens, effects, lookback, lens, ragas, selfcheckgpt, evaluator, longfact, safe, factscore, truthfulness, truthfulqa, speed, 220k, ssms, faster, hybrids, balance, efficiency, characterizing, undertraining, frequency, dimensions, evolutionary, rescaling, llama3, 80x, fewer, longrope2, coarse, compression, selection, speedups, 64k, nsa, 456b, lightning, softmax, minimax, module, scales, outperforms, titans, extending, interpolation, jamba, rnn, counts, rwkv, selective, space, 1000, controlled, experiments, recipes, closed, openthoughts, prms, supervision, thinkprm, gae, stable, vapo, key, decoupled, entropy, bonus, reproduces, dapo, paired, rollouts, bootstrap, distilling, rstar, prm, reaching, kimi, budget, forcing, optimally, replicated, let, verify, reasoners, reflexion, refine, trees, thoughts, majority, vote, sampled, chains, result, intermediate, steps, trend, defining, traces, benefit, extra, configurations, mitigates, degrades, niah, 503, expert, crafted, spanning, humans, pressure, ruler, probe, 824, requiring, frames, conversational, simultaneous, tested, frontier, below, multichallenge, typologically, prox, harness, contamination, resistant, monthly, livebench, elo, leaderboard, arena, lmsys, verifiable, ifeval, scientific, overclaim, missing, refute, graduate, proof, gpqa, harder, successor, multitask, subjects, capabilities, bench, holistic, helm, heterogeneous, massive, xglue, superglue, glue, xiaomi, scaled, commercial, milmmt, specialized, translategemma, mcgill, 14b, continued, 26b, empirical, mixing, afriquellm, princeton, mila, adapted, wura, irokobench, afriqa, lugha, 83b, population, speakers, comparably, sized, xcopa, mgsm, babel, targeting, seallm, seamless, left, glot500, expanse, 176b, bloom, mt5, commoncrawl, xlm, objective, generalization, mixtral, mid, reproducibility, olmo, especially, often, framing, 250m, depth, width, identical, neobert, modernized, flashattention, replaced, sample, electra, disentangled, robustly, bidirectional, workhorse, them, scoped, phenomena, mostly, sami, mordvin, mari, komi, supported, finnish, swedish, splitting, dynet, standalone, cli, bindings, rest, cube, tokenizing, lemmatizing, primarily, solution, tiny, own, copies, tiny_qa_benchmark_pp, mixture, 167, culturax, 15t, cleaned, filters, educational, documented, filtering, dolma, reproductions, 30t, signals, 825, gib, pile, central, versioned, streamable, hubs, coqui, wav2vec, 170, realtime, gpu, vad, punctuation, diarization, emotion, autoregressive, sensevoice, fun, nano, funasr, nvidia, canary, whisper, borders, expansion, marco, lotte, att, intensive, remixer, synthesis, redapter, record, ndcg, bright, reasonembed, reranking, ood, rank1, surpassing, prior, proprietary, derived, xor, original, variable, dimensionality, matryoshka, representation, embed, families, colbertv2, dual, passage, increasingly, novel, rate, schedule, bea, edit, gec, surpass, role, docred, privee, automatically, policies, templates, schema, openie, assistant, gaia, fid, drqa, distantly, hotpotqa, have, restructured, bridging, divide, sub, 10b, turbo, comparison, professional, translators, chatgpt, translator, similarity, bertscore, marian, reset, field, fluency, hurts, grounding, longer, budgets, harm, element, summarizers, benchmarking, scrolls, booksum, pegasus, graph, textrank, anchor, corex, jointly, top2vec, bertopic, lsi, hdp, blei, caveats, annotators, imdb, lead, closing, shared, dethrone, motivated, lingmess, maverick, order, hoi, spanbert, lee, eight, replace, guideline, gollie, generalist, arbitrary, gliner, string, bilstm, lample, 2003, attentive, kitaev, klein, light, trankit, biaffine, hoc, additions, coalesce, sequences, retraining, premiums, quantifies, fertility, predicts, penalties, morphologically, tax, iclr, formal, stochastic, map, establishes, conditions, foundations, decouples, output, vocabularies, log, relationship, independently, superword, downstream, superbpe, patching, revives, blt, operating, characters, canine, byt5, wordpiece, units, pair, encoding, sennrich, morfessor, two, dominant, schemes, infersent, cove, doc2vec, sense2vec, handle, oov, static, problem, each, subsection, linking, mace, checks, assisted, 300, example, potato, modal, studio, collecting, curating, rich, assertion, unlimited, documents, foss
Text of the page (random words):
hoonya is data agnostic can be used by teams to annotate data with various level of verification stages at scale annotation lab free end to end no code platform for text annotation and dl model training tuning out of the box support for named entity recognition classification relation extraction and assertion status spark nlp models unlimited support for users teams projects documents not foss flat flat is a web based linguistic annotation environment based around the folia format a rich xml based format for linguistic annotation free and open source argilla open source platform for collecting human feedback building nlp and llm datasets and curating preference data label studio open core multi modal labeling platform widely used for nlp labeling potato free open source annotation tool covering 21 task types classification span coreference entity linking agent trace evaluation with built in mace quality control attention checks ai assisted labeling and 300 example tasks tasks and methods nlp tasks organized by linguistic problem each subsection lists foundational classical work first then neural approaches then llm based methods where relevant for modern lm specific research pretraining evaluation retrieval reasoning etc see language models for nlp text embeddings back to top static word embeddings foundational word2vec implementation explainer blog glove explainer blog fasttext implementation subword n grams handle oov well still useful for low resource languages sense2vec word sense disambiguation paragraph vectors doc2vec contextual embeddings elmo deep contextualized word representations cove contextualized vectors learned from mt ulmfit language model fine tuning for text classification infersent sentence representations from nli modern sentence and document embeddings see retrieval for nlp sentence transformers e5 bge m3 nomic gritlm and mteb for current leaderboards tokenization morphology and segmentation back to top sentencepiece language agnostic subword tokenization bpe and unigram lm the two dominant subword schemes stanza tokenization lemma and morphology for 70 languages udpipe tokenization tagging lemmatization parsing for universal dependencies morfessor unsupervised morphological segmentation tokenizer research and architecture also see language models byte pair encoding sennrich et al subword units for neural mt foundation of modern tokenizers sentencepiece language agnostic subword tokenization bpe and unigram tokenizers fast rust implementations of bpe wordpiece unigram byt5 tokenizer free byte level model canine tokenization free encoder operating on unicode characters how good is your tokenizer tokenizer fairness across languages byte latent transformer blt meta 2024 dynamic byte level patching that matches bpe tokenized models at scale revives the tokenizer free direction superbpe 2025 superword tokenization that improves on bpe for downstream tasks over tokenized transformer icml 2025 decouples input and output vocabularies shows a log linear relationship between input vocabulary size and training loss scaling vocabulary independently of model size foundations of tokenization iclr 2025 first formal unified framework for tokenizer models using stochastic map category theory establishes conditions for statistical consistency the token tax systematic bias in multilingual tokenization 2025 quantifies how tokenization fertility predicts model accuracy across languages exposing structural cost penalties for morphologically complex and low resource languages reducing tokenization premiums for low resource languages 2026 post hoc vocabulary additions that coalesce multi token character sequences for low resource languages reducing inference cost without retraining pos tagging and dependency parsing back to top universal dependencies cross linguistically consistent treebanks 100 languages spacy and stanza production parsers across many languages deep biaffine attention for neural dependency parsing foundational neural parsing architecture trankit light weight transformer based multilingual nlp toolkit self attentive constituency parsing kitaev klein strong neural constituency parser named entity recognition and information extraction back to top foundational and neural conll 2003 ner canonical english ner benchmark neural architectures for ner lample et al bilstm crf the long time go to ner architecture flair contextual string embeddings strong ner across languages spacy ner production ready open and instruction following ie universal ner instruction tuned lm for open set ner across languages gliner 2023 small generalist ner model that handles arbitrary entity types at inference gollie guideline following information extraction with lms rebel end to end relation extraction as seq2seq llm based gpt ner llms for named entity recognition can llms replace sentence level ner 2024 cost quality tradeoffs generative ner in the era of llms 2026 eight open llms across four ner benchmarks peft with structured outputs matches encoder based ner coreference resolution back to top end to end neural coreference lee et al foundation for modern neural coreference spanbert span based pretraining strong coreference baseline coref hoi higher order inference coreference maverick coref 2024 efficient coreference matching the best larger systems lingmess linguistically motivated category based coreference scoring llm based llms for coreference resolution prompting and fine tuning for coreference multilingual coreference shared task can llms dethrone traditional approaches 2025 9 systems across 4 llm based and 5 traditional approaches traditional methods still lead but llms are closing the gap text classification and sentiment analysis back to top fasttext classifier strong fast linear baseline sentiment treebank sst canonical fine grained sentiment dataset setfit few shot text classification without prompts fastfit fast few shot for many class settings sst imdb ag news with deberta v3 current encoder fine tuning baseline pyss3 white box interpretable text classifier llms as annotators using llms for text classification labeling with caveats topic modeling back to top latent dirichlet allocation blei et al foundational topic model gensim lda lsi hdp in python bigartm fast regularized topic modeling bertopic clustering based topic modeling on top of contextual embeddings common modern default top2vec jointly learns topic and document vectors corex topic hierarchical topic modeling with anchor words summarization back to top textrank extractive graph based summarization pointer generator networks see et al foundational neural abstractive summarization pegasus gap sentences pretraining for summarization bart widely used denoising seq2seq baseline booksum and scrolls long document summarization benchmarks llm based benchmarking llms for news summarization llms vs fine tuned summarizers element aware summarization with llms structured prompting for summarization understanding llm reasoning for abstractive summarization 2025 explicit reasoning improves fluency but hurts factual grounding longer reasoning budgets can harm faithfulness machine translation back to top statistical and foundational neural moses reference statistical mt system attention is all you need transformer reset the field marian nmt efficient c nmt framework fairseq pytorch sequence modeling toolkit massively multilingual nllb 200 mt for 200 languages madlad 400 400 language mt seamlessm4t speech and text mt 100 languages evaluation comet learned mt metric current de facto standard alongside chrf sacrebleu reproducible bleu chrf ter scoring bertscore similarity based generation metric llm based is chatgpt a good translator llms as machine translation systems adapting llms for document level mt 2024 llms for context aware translation gpt 4 vs human translators quality comparison on professional mt multilingual mt with open llms at practical scale 2025 benchmarks sub 10b open llms on 28 language mt matches gpt 4 turbo and google translate bridging the linguistic divide survey on llms for mt 2025 survey of how instruction following in context learning and preference alignment have restructured mt methodology question answering and reading comprehension back to top datasets and foundational systems squad squad 2 0 extractive reading comprehension natural questions real user questions over wikipedia hotpotqa multi hop reasoning triviaqa distantly supervised qa drqa open domain qa over wikipedia document qa multi paragraph reading comprehension modern open domain qa dpr and fid retrieve then read the standard pre llm open domain qa pipeline atlas retrieval augmented lm for few shot qa see also retrieval for nlp llm era gpt 4 with retrieval on triviaqa nq self rag 2023 retrieval generation and self critique gaia general ai assistant benchmark including multi step qa information extraction beyond ner back to top openie 6 schema free open information extraction template based information extraction without the templates privee an architecture for automatically analyzing web privacy policies rebel end to end relation extraction docred document level relation extraction benchmark llms for semantic role labeling 2025 generative llms with rag and self correction surpass encoder decoder bert style models on srl in english and chinese adapting llms for minimal edit gec 2025 decoder only llms with a novel error rate adaptation schedule set new sota on bea test grammatical error correction retrieval and embeddings back to top dense and late interaction retrieval increasingly the substrate for qa and ir dpr dense passage retrieval dual encoder retrieval baseline colbert and colbertv2 late interaction retrieval strong on out of domain e5 and e5 mistral widely used dense embedding families bge and bge m3 2024 multilingual multi functionality embeddings top of mteb across languages nomic embed 2024 fully open reproducible embedding model matryoshka representation learning nested embeddings supporting variable dimensionality at inference gritlm 2024 unified generation and embedding from one model rag retrieval augmented generation the original retrieval augmented framework foundation for modern qa pipelines gemini embedding 2025 gemini derived dense embeddings sota on mmteb across 250 languages and on cross lingual retrieval xor retrieve xtreme up qwen3 embedding 2025 decoder based embedding series 0 6b 8b built on qwen3 1 on mteb multilingual and mteb code surpassing prior proprietary models rank1 2025 first reranking model trained with test time compute via deepseek r1 reasoning trace distillation sota on instruction following and ood retrieval reasonembed 2025 embedding model for reasoning intensive retrieval with remixer data synthesis and redapter adaptive training record ndcg 10 of 38 1 on bright colbert att 2026 extends late interaction retrieval by integrating query and document attention weights into colbert scoring improves recall on ms marco beir and lotte embedding and retrieval benchmarks mmteb 2025 community expansion of mteb to 500 tasks across 250 languages speech and text back to top a short pointer set since this borders adjacent fields whisper multilingual asr the modern open default seamlessm4t unified speech and text translation canary nvidia 2024 top open multilingual asr model funasr industrial grade asr toolkit 170 realtime on gpu 50 languages built in vad punctuation speaker diarization and emotion detection includes non autoregressive sensevoice and llm based fun asr nano models wav2vec 2 0 foundational self supervised speech pretraining coqui tts and vieneu tts open tts datasets back to top dataset hubs and lists huggingface datasets hub the central index for modern nlp datasets with versioned streamable loaders nlp datasets large collection of nlp datasets gensim data data repository for pretrained nlp models and nlp corpora pretraining scale corpora open the pile 825 gib diverse text corpus redpajama redpajama v2 2023 2024 reproductions of llama pretraining data v2 is 30t tokens with quality signals dolma ai2 2023 2024 3t token open pretraining corpus with documented filtering pipeline fineweb fineweb edu 2024 15t token cleaned web corpus fineweb edu filters for educational quality culturax 6 3t tokens across 167 languages common corpus 2024 2t token open license multilingual corpus task and instruction datasets universal dependencies cross linguistically consistent treebank annotation 100 languages tülu 3 sft mixture 2024 open instruction tuning data behind tülu 3 tiny_qa_benchmark_pp tiny nlp multi lingual qa datasets and library to generate your own synthetic copies multilingual nlp frameworks back to top udpipe is a trainable pipeline for tokenizing tagging lemmatizing and parsing universal treebanks and other conll u files primarily written in c offers a fast and reliable solution for multilingual nlp processing nlp cube natural language processing pipeline sentence splitting tokenization lemmatization part of speech tagging and dependency parsing new platform written in python with dynet 2 0 offers standalone cli python bindings and server functionality rest api uralicnlp is an nlp library mostly for many endangered uralic languages such as sami languages mordvin languages mari languages komi languages and so on also some non endangered languages are supported such as finnish together with non uralic languages such as swedish and arabic uralicnlp can do morphological analysis generation lemmatization and disambiguation language models for nlp back to top pretrained language models and the research around them scoped to nlp tasks and linguistic phenomena for general purpose llm tooling agents or rag application kits see see also pretraining and adaptation encoders still the workhorse for classical nlp tasks bert bidirectional transformer pretraining foundation for most encoder based nlp work since 2018 roberta robustly optimized bert pretraining common encoder baseline deberta deberta v3 disentangled attention strong on classification ner nli electra replaced token detection pretraining sample efficient modernbert 2024 modernized encoder with rotary embeddings flashattention 8k context current go to encoder for classification ner retrieval neobert 2025 250m parameter encoder integrating modern architecture improvements rope 4k context optimized depth to width state of the art on mteb surpasses modernbert and roberta large under identical fine tuning encoder decoder and seq2seq t5 and flan t5 text to text framing for nlp tasks strong instruction tuned encoder decoder baselines bart denoising seq2seq pretraining widely used for summarization and generation open decoder only lms used as substrate for nlp tasks llama 3 3 1 3 3 meta 2024 2025 widely adopted open weight family default base for fine tuning across nlp tasks qwen 2 5 qwen 3 alibaba 2024 2025 strong multilingual coverage especially chinese often top open model on multilingual benchmarks de...
|