Meta tags:
description= :book: A curated list of resources dedicated to Natural Language Processing (NLP) - keon/awesome-nlp;
Headings (most frequently used words):
and, nlp, in, models, libraries, embeddings, datasets, language, data, uh, oh, navigation, text, saved, searches, awesome, files, research, tutorials, multilingual, license, footer, reading, information, extraction, languages, corpora, search, code, repositories, users, issues, pull, requests, provide, feedback, keon, menu, use, to, filter, your, results, more, quickly, folders, latest, commit, history, repository, scope, contents, summaries, trends, prominent, labs, tasks, methods, frameworks, for, per, see, also, citation, about, releases, packages, contributors, historical, highlights, content, videos, online, courses, books, services, annotation, tools, tokenization, morphology, segmentation, pos, tagging, dependency, parsing, named, entity, recognition, coreference, resolution, classification, sentiment, analysis, topic, modeling, summarization, machine, translation, question, answering, comprehension, beyond, ner, retrieval, speech, pretraining, adaptation, cross, lingual, evaluation, benchmarks, reasoning, test, time, compute, long, context, alternative, architectures, factuality, hallucination, calibration, probing, interpretability, efficient, small, instruction, tuning, preference, optimization, bias, fairness, safety, arabic, chinese, anthology, danish, dutch, german, hungarian, indic, treebanks, word, tooling, indonesian, korean, blogs, persian, polish, portuguese, spanish, thai, ukrainian, urdu, uzbek, vietnamese, other, topics, resources, contributing, stars, watchers, forks, that, need, login, access, can, be, gained, via, email,
Text of the page (most frequently used words):
and (360), for (250), nlp (233), #language (124), with (108), the (103), models (100), open (78), languages (71), natural (65), 2024 (64), text (63), top (63), processing (61), 2025 (61), back (54), based (47), library (44), multilingual (43), model (42), data (40), python (37), from (36), embeddings (35), across (33), ner (33), tasks (33), datasets (32), word (32), learning (30), llms (29), reasoning (28), retrieval (27), toolkit (27), pos (26), annotation (26), korean (25), evaluation (25), benchmark (25), level (25), corpus (24), free (24), long (24), libraries (23), tagging (23), persian (23), bert (22), context (22), neural (22), llm (21), speech (21), machine (20), chinese (20), instruction (20), tokenization (20), lms (20), extraction (20), research (20), arabic (20), resources (19), deep (19), parsing (19), classification (19), analysis (19), code (19), generation (18), large (18), vietnamese (18), dependency (18), translation (18), multi (18), training (18), topic (18), github (17), awesome (17), tools (17), sentiment (17), hindi (17), pretraining (17), english (16), treebank (16), 2023 (16), spanish (16), encoder (16), source (16), 2026 (15), list (15), summarization (15), transformer (15), modeling (15), that (14), time (14), corpora (14), covering (14), sentence (14), strong (14), trained (14), embedding (14), coreference (14), document (14), tool (14), curated (13), rag (13), see (13), dataset (13), package (13), efficient (13), platform (13), small (13), tuning (13), transformers (13), attention (13), modern (13), information (12), also (12), linguistic (12), indonesian (12), tokens (12), scale (12), fast (12), recognition (12), entity (12), cross (12), fine (12), search (12), token (12), test (12), all (12), your (12), end (12), this (11), including (11), task (11), benchmarks (11), indic (11), quality (11), parser (11), named (11), inference (11), built (11), you (10), can (10), security (10), production (10), collection (10), thai (10), implementation (10), other (10), family (10), pretrained (10), segmentation (10), foundation (10), word2vec (10), course (10), high (10), spacy (10), low (10), state (10), features (10), view (10), foundational (10), such (10), systems (9), tokenizer (9), bilingual (9), sentences (9), universal (9), tuned (9), java (9), methods (9), morphological (9), pytorch (9), dutch (9), support (9), over (9), resource (9), framework (9), bias (9), more (9), art (9), self (9), decoder (9), sparse (9), hallucination (9), like (9), labeling (9), license (8), general (8), current (8), using (8), standard (8), first (8), format (8), llama (8), university (8), via (8), common (8), human (8), shot (8), mteb (8), google (8), web (8), are (8), performance (8), architecture (8), 200 (8), lingual (8), api (8), error (7), please (7), reload (7), asian (7), words (7), semantic (7), sequence (7), urdu (7), portuguese (7), used (7), analyzer (7), scala (7), nltk (7), preference (7), feedback (7), huggingface (7), aware (7), dense (7), unified (7), survey (7), factuality (7), calibration (7), detection (7), question (7), answering (7), compute (7), stanford (7), community (6), contributing (6), topics (6), prompt (6), generative (6), here (6), 100 (6), news (6), southeast (6), relation (6), available (6), part (6), fully (6), dependencies (6), suite (6), samples (6), pipeline (6), danish (6), moe (6), mmlu (6), domain (6), alignment (6), safety (6), matches (6), widely (6), written (6), without (6), interpretable (6), representations (6), probing (6), gpt (6), linear (6), sota (6), system (6), following (6), nlu (6), seq2seq (6), baseline (6), reading (6), how (6), tokenizers (6), building (6), teams (6), various (6), clojure (6), node (6), enterprise (6), navigation (5), there (5), was (5), while (5), repository (5), keon (5), serving (5), prompting (5), augmented (5), papers (5), classical (5), japanese (5), parallel (5), core (5), 400 (5), fasttext (5), covers (5), roberta (5), only (5), wikipedia (5), wiki (5), advanced (5), blog (5), tutorials (5), browser (5), alternative (5), scratch (5), elmo (5), iit (5), german (5), lemmatization (5), deepseek (5), book (5), per (5), many (5), resolution (5), sft (5), reference (5), optimization (5), parameter (5), structured (5), few (5), settings (5), memory (5), larger (5), distillation (5), autoencoders (5), interpretability (5), scaling (5), dynamic (5), pipelines (5), hybrid (5), meta (5), beyond (5), simple (5), loaders (5), statistical (5), traditional (5), ready (5), bpe (5), subword (5), folia (5), spark (5), projects (5), use (5), online (5), actions (5), content (5), include (5), network (5), creating (5), kotlin (5), group (5), notebooks (5), issues (5), summaries (5), illustrated (5), loading (4), page (4), readme (4), mining (4), dedicated (4), find (4), tooling (4), purpose (4), lists (4), out (4), emnlp (4), tts (4), ukrainian (4), best (4), polish (4), annotated (4), most (4), tags (4), nmt (4), technology (4), style (4), provides (4), polarity (4), facebook (4), review (4), stopwords (4), hungarian (4), focus (4), qwen (4), correction (4), morphology (4), anthropic (4), fairness (4), results (4), zero (4), weight (4), adaptation (4), layer (4), cache (4), quantization (4), bge (4), paragraph (4), accuracy (4), non (4), gemma (4), reproducible (4), than (4), clustering (4), architectures (4), feature (4), transcoders (4), outputs (4), hop (4), knowledge (4), factual (4), form (4), set (4), grained (4), extension (4), matching (4), process (4), step (4), surpasses (4), into (4), questions (4), refresh (4), translate (4), behind (4), work (4), size (4), new (4), functionality (4), files (4), grade (4), asr (4), includes (4), scoring (4), series (4), then (4), pre (4), comprehension (4), networks (4), vectors (4), canonical (4), approaches (4), byte (4), rust (4), wide (4), supports (4), costs (4), powered (4), interface (4), services (4), julia (4), ruby (4), regular (4), metrics (4), visual (4), notable (4), acl (4), pull (4), action (3), terms (3), report (3), stars (3), cc0 (3), about (3), 2018 (3), useful (3), citation (3), techniques (3), template (3), engineering (3), hebrew (3), texts (3), elasticsearch (3), russian (3), good (3), tagger (3), books (3), 250 (3), big (3), constituency (3), sailor (3), mistral (3), chat (3), mit (3), uzbek (3), focused (3), etc (3), multiple (3), center (3), both (3), annotations (3), tagged (3), stemmer (3), squad (3), archive (3), movie (3), institute (3), bahasa (3), collections (3), indobert (3), telugu (3), patna (3), access (3), labelled (3), sense (3), twitter (3), domains (3), aspect (3), bangla (3), classes (3), treebanks (3), developed (3), industrial (3), binding (3), frog (3), alibaba (3), trainable (3), javascript (3), faking (3), policy (3), value (3), reveals (3), jailbreak (3), integrating (3), turn (3), strategies (3), pairs (3), grpo (3), way (3), ai2 (3), post (3), recipe (3), tülu (3), adapting (3), class (3), loss (3), qwen3 (3), 128k (3), competitive (3), much (3), encoders (3), specific (3), sae (3), shows (3), introduces (3), graphs (3), tracing (3), methodology (3), recall (3), learns (3), claim (3), uncertainty (3), before (3), through (3), faithfulness (3), under (3), comprehensive (3), rope (3), extends (3), recurrent (3), tradeoffs (3), patterns (3), minimal (3), mamba (3), reward (3), produce (3), chain (3), thought (3), adaptive (3), math (3), cot (3), short (3), supervised (3), needle (3), haystack (3), but (3), together (3), pro (3), gap (3), between (3), gemini (3), seamlessm4t (3), 101 (3), transfer (3), default (3), optimized (3), deberta (3), still (3), application (3), endangered (3), disambiguation (3), frameworks (3), linguistically (3), consistent (3), fineweb (3), gensim (3), fields (3), late (3), interaction (3), colbert (3), improves (3), supporting (3), dpr (3), analyzing (3), read (3), real (3), practical (3), bleu (3), chrf (3), learned (3), moses (3), contextual (3), latent (3), box (3), classifier (3), higher (3), cost (3), types (3), flair (3), stanza (3), vocabulary (3), input (3), unicode (3), implementations (3), unigram (3), agnostic (3), well (3), glove (3), where (3), environment (3), xml (3), shoonya (3), cloud (3), intent (3), syntax (3), working (3), variety (3), engine (3), archived (3), snips (3), software (3), opennlp (3), exploring (3), developing (3), applications (3), expression (3), forms (3), jupyter (3), lectures (3), recent (3), labs (3), penn (3), trends (3), repositories (3), commit (3), name (3), insights (3), requests (3), signed (3), another (3), tab (3), window (3), session (3), sign (3), saved (3), documentation (3), copilot (3), solutions (3), explore (3), perform (2), not (2), manage (2), status (2), privacy (2), footer (2), releases (2), forks (2), adjacent (2), scope (2), ancient (2), 10k (2), device (2), vieneu (2), grounded (2), speeches (2), prime (2), minister (2), tensorflow (2), character (2), coverage (2), national (2), reliable (2), uppsala (2), translated (2), 250k (2), suitable (2), french (2), petition (2), chatbot (2), naver (2), blogs (2), klue (2), native (2), mecab (2), multimodal (2), hub (2), providing (2), standardized (2), wordnet (2), sea (2), lion (2), cendol (2), indolem (2), indonlu (2), algorithm (2), sarvam (2), sanskrit (2), bengali (2), whole (2), ulmfit (2), marathi (2), need (2), email (2), labels (2), absa (2), reviews (2), smaller (2), strength (2), kaldi (2), anthology (2), jieba (2), released (2), alongside (2), arabert (2), organized (2), cases (2), when (2), vector (2), mitigation (2), refusal (2), attacks (2), failures (2), user (2), measurement (2), gender (2), masked (2), instruct (2), rlhf (2), less (2), adopted (2), direct (2), instructions (2), follow (2), flan (2), lora (2), peft (2), rank (2), quantized (2), throughput (2), icml (2), precision (2), relevant (2), fastfit (2), prompts (2), setfit (2), apple (2), intelligence (2), thinking (2), technical (2), local (2), global (2), yarn (2), microsoft (2), ones (2), phi (2), distilbert (2), identifies (2), circuits (2), highly (2), inputs (2), skip (2), applies (2), attribution (2), case (2), enables (2), circuit (2), causal (2), computational (2), examples (2), behavior (2), browsing (2), extracting (2), monosemanticity (2), structural (2), what (2), taxonomy (2), choice (2), pattern (2), sampling (2), ssm (2), combining (2), degradation (2), contexts (2), term (2), memorize (2), historical (2), lost (2), middle (2), windows (2), position (2), rotary (2), length (2), aime (2), baselines (2), verification (2), discriminative (2), clipping (2), four (2), improvements (2), mcts (2), systematic (2), study (2), openai (2), pure (2), tree (2), consistency (2), improve (2), direction (2), explicit (2), nested (2), score (2), longbench (2), synthetic (2), exposing (2), diverse (2), evidence (2), critique (2), beir (2), flores (2), xnli (2), xtreme (2), african (2), speaker (2), madlad (2), nllb (2), 500 (2), cohere (2), massively (2), aya (2), weights (2), base (2), substrate (2), denoising (2), bart (2), modernbert (2), nli (2), since (2), around (2), agents (2), kits (2), uralic (2), some (2), uralicnlp (2), offers (2), server (2), conll (2), udpipe (2), generate (2), edu (2), redpajama (2), index (2), pointer (2), mmteb (2), query (2), trace (2), retrieve (2), one (2), gritlm (2), nomic (2), grammatical (2), srl (2), rebel (2), triviaqa (2), era (2), atlas (2), extractive (2), metric (2), ter (2), sacrebleu (2), facto (2), comet (2), fairseq (2), understanding (2), abstractive (2), generator (2), hierarchical (2), regularized (2), bigartm (2), lda (2), dirichlet (2), allocation (2), white (2), pyss3 (2), sst (2), category (2), coref (2), span (2), handles (2), crf (2), parsers (2), reducing (2), complex (2), theory (2), tokenized (2), sentencepiece (2), unsupervised (2), lemma (2), leaderboards (2), contextualized (2), grams (2), explainer (2), agent (2), control (2), label (2), argilla (2), flat (2), users (2), lab (2), management (2), easy (2), hosted (2), active (2), freemium (2), collaborative (2), discourse (2), offering (2), doccano (2), brat (2), rapid (2), custom (2), cognitive (2), service (2), amazon (2), related (2), vscode (2), clojurescript (2), visualizations (2), mallet (2), apache (2), epic (2), alike (2), ucto (2), bllip (2), crfsuite (2), conditional (2), random (2), crfs (2), sequential (2), evaluate (2), dsl (2), allows (2), which (2), regex (2), industry (2), build (2), very (2), extensible (2), modelling (2), jptdp (2), interactive (2), better (2), plain (2), adversarial (2), world (2), beginner (2), friendly (2), hands (2), hugging (2), face (2), authors (2), program (2), rnns (2), taking (2), introduction (2), carnegie (2), mellon (2), oxford (2), videos (2), courses (2), paper (2), guides (2), explanations (2), guide (2), sebastian (2), ruder (2), newsletter (2), playbook (2), engineer (2), their (2), contributions (2), development (2), create (2), project (2), prominent (2), rolling (2), venues (2), they (2), add (2), pull_request_template (2), assets (2), 622 (2), commits (2), last (2), message (2), menu (2), appearance (2), cancel (2), searches (2), business (2), developer (2), customer (2), devops (2), app (2), share, personal, cookies, contact, docs, inc, contributors, packages, published, watching, 568, watchers, activity, note, https, com, url, year, kim, woo, author, title, misc, consider, citing, mlops, modalities, nlph_resources, doing, cltk, lao, icu, pymorphy2, 20m, law, articles, film, subtitles, evb, sql, 2020, findings, vitext2sql, 75m, vntqcorpus, txt, hours, recorded, hcmus, ailab, vivos, ud_vietnamese, bktreebank, vistral, vinai, phogpt, bartpho, phobert, voice, cloning, pyvi, vncorenlp, vitk, underthesea, culturally, rows, soas, urduhack, ukrainianlt, thailand, inter, openthaigpt, scb, 10x, typhoon, wangchanberta, synthai, cutkum, cluster, jtcc, pythainlp, sent2vec, rigochat, bsc, barcelona, supercomputing, salamandra, basque, latxa, bne, beto, compilation, unannotated, billion, copenhagen, columbian, political, detect, censor, clean, profanity, hate, bullying, speaking, countries, spanlp, portulan, albertina, maritaca, sabiá, brazilian, bertimbau, clef, 2008, 2009, hamshahri, syntactically, updt, 982, verbs, valency, lexicon, syntactic, semeval, 2010, perlex, 25m, farsiyar, persianner, 682, iob, armanpersonercorpus, lscp, 120m, 27m, casual, tweets, colloquial, freely, upc, farsi, manually, bijankhan, dorna, persianmind, parsbert, cleaning, virastar, parsianalyzer, partial, perstem, keyphrase, perke, parsivar, hazm, html, korquad, expired, blue, house, site, petitions, major, south, newspaper, chosun, ilbo, korea, science, kaist, kangwon, dsindex, hyperclova, exaone, polyglot, skt, kobert, running, client, side, webassembly, 1mb, offline, garu, kiwi, splitter, kss, konlp, koalanlp, konlpy, seacrowd, dictionary, indosum, idn, 39k, 900k, panl10n, kompas, tempo, ilps, singapore, regional, nusacrowd, indonesia, sastrawi, stemming, pysastrawi, bharatgpt, krutrim, airavata, continuation, openhathi, indictrans2, 2022, indicbert, ai4bharat, indicnlp, fastai, inltk, kannada, sivareddy, python3, port, transliteration, helpers, oscar, albert, bunch, crawl, languge, hindi2vec, tdil, aggregates, lot, otherwise, gated, sentiwordnet, bombay, cfilt, tamil, sail, 2015, login, gained, bbc, 60k, peter, graham, isi, fire, above, mentioned, representational, layered, off, shelf, does, alpino, surface, realiser, simplenlg, simplenlg_nl, danlp, funnlp, baichuan, tsinghua, chatglm3, glm, improved, mlm, macbert, masking, wwm, hit, ltp, hanlp, fudannlp, snownlp, arabicmmlu, aggregated, labr, largest, multidomain, sdaia, allam, jais, araelectra, msa, dialectal, camelbert, qcri, farasa, segmenter, coptic, rftokenizer, pyarabic, jsastem, goarabic, dialect, camel, click, section, expand, occurs, conflicts, internalized, values, steering, reduces, vlaf, conflict, agreement, scripts, indicsafe, modular, defenses, risk, categories, teleai, 4000, dialogues, scenarios, safedialbench, finetuning, narrow, insecure, unexpectedly, produces, broad, unrelated, emergent, misalignment, moderation, wildguard, strategically, complying, during, tailoring, answers, beliefs, sycophancy, toxicity, realtoxicityprompts, demographic, axes, holisticbias, winobias, social, crows, measuring, stereotypical, stereoset, synthesizes, response, aligned, nothing, filtered, subset, official, magpie, dpo, trl, goes, lima, among, simpler, generated, against, constitution, constitutional, 1600, super, naturalinstructions, bootstrapping, instructgpt, finetuned, learners, bundling, prefix, ia3, others, decomposed, dora, adapters, modest, hardware, qlora, tgi, sglang, pagedattention, vllm, portable, gguf, cpp, sensitivity, wise, mixed, improvement, uniform, kv8, kvtuner, activation, awq, gptq, deploying, compact, near, stella, gte, siamese, sharing, bit, qat, reduction, 235b, modes, 30b, a3b, activating, parameters, 27b, ratio, keep, tractable, nope, smollm3, smollm2, distilled, minilm, rather, mechanism, finding, explanation, reconstructing, yield, saes, beat, claude, haiku, rhyme, planning, studies, biology, construct, replacement, interactions, revealing, identifying, driving, influence, functions, neuronpedia, towards, toy, superposition, pyramid, probes, locating, editing, associations, rome, limitations, alternatives, classifiers, belinkov, primer, bertology, trains, reason, generating, gains, biography, factbench, auroc, cure, think, hard, required, responses, rates, persist, even, halluhard, logits, principled, quantification, fact, checking, formally, separates, franq, substantially, worse, calibrated, extended, single, claims, atomic, extrinsic, intrinsic, regeneration, resist, leakage, hallulens, effects, lookback, lens, ragas, selfcheckgpt, evaluator, longfact, safe, factscore, truthfulness, truthfulqa, speed, 220k, ssms, faster, hybrids, balance, efficiency, characterizing, undertraining, frequency, dimensions, evolutionary, rescaling, llama3, 80x, fewer, longrope2, coarse, compression, selection, speedups, 64k, nsa, 456b, lightning, softmax, minimax, module, scales, outperforms, titans, extending, interpolation, jamba, rnn, counts, rwkv, selective, space, 1000, controlled, experiments, recipes, closed, openthoughts, prms, supervision, thinkprm, gae, stable, vapo, key, decoupled, entropy, bonus, reproduces, dapo, paired, rollouts, bootstrap, distilling, rstar, prm, reaching, kimi, budget, forcing, optimally, replicated, let, verify, reasoners, reflexion, refine, trees, thoughts, majority, vote, sampled, chains, result, intermediate, steps, trend, defining, traces, benefit, extra, configurations, mitigates, degrades, niah, 503, expert, crafted, spanning, humans, pressure, ruler, probe, 824, requiring, frames, conversational, simultaneous, tested, frontier, below, multichallenge, typologically, prox, harness, contamination, resistant, monthly, livebench, elo, leaderboard, arena, lmsys, verifiable, ifeval, scientific, overclaim, missing, refute, graduate, proof, gpqa, harder, successor, multitask, subjects, capabilities, bench, holistic, helm, heterogeneous, massive, xglue, superglue, glue, xiaomi, scaled, commercial, milmmt, specialized, translategemma, mcgill, 14b, continued, 26b, empirical, mixing, afriquellm, princeton, mila, adapted, wura, irokobench, afriqa, lugha, 83b, population, speakers, comparably, sized, xcopa, mgsm, babel, targeting, seallm, seamless, left, glot500, expanse, 176b, bloom, mt5, commoncrawl, xlm, objective, generalization, mixtral, mid, reproducibility, olmo, especially, often, framing, 250m, depth, width, identical, neobert, modernized, flashattention, replaced, sample, electra, disentangled, robustly, bidirectional, workhorse, them, scoped, phenomena, mostly, sami, mordvin, mari, komi, supported, finnish, swedish, splitting, dynet, standalone, cli, bindings, rest, cube, tokenizing, lemmatizing, primarily, solution, tiny, own, copies, tiny_qa_benchmark_pp, mixture, 167, culturax, 15t, cleaned, filters, educational, documented, filtering, dolma, reproductions, 30t, signals, 825, gib, pile, central, versioned, streamable, hubs, coqui, wav2vec, 170, realtime, gpu, vad, punctuation, diarization, emotion, autoregressive, sensevoice, fun, nano, funasr, nvidia, canary, whisper, borders, expansion, marco, lotte, att, intensive, remixer, synthesis, redapter, record, ndcg, bright, reasonembed, reranking, ood, rank1, surpassing, prior, proprietary, derived, xor, original, variable, dimensionality, matryoshka, representation, embed, families, colbertv2, dual, passage, increasingly, novel, rate, schedule, bea, edit, gec, surpass, role, docred, privee, automatically, policies, templates, schema, openie, assistant, gaia, fid, drqa, distantly, hotpotqa, have, restructured, bridging, divide, sub, 10b, turbo, comparison, professional, translators, chatgpt, translator, similarity, bertscore, marian, reset, field, fluency, hurts, grounding, longer, budgets, harm, element, summarizers, benchmarking, scrolls, booksum, pegasus, graph, textrank, anchor, corex, jointly, top2vec, bertopic, lsi, hdp, blei, caveats, annotators, imdb, lead, closing, shared, dethrone, motivated, lingmess, maverick, order, hoi, spanbert, lee, eight, replace, guideline, gollie, generalist, arbitrary, gliner, string, bilstm, lample, 2003, attentive, kitaev, klein, light, trankit, biaffine, hoc, additions, coalesce, sequences, retraining, premiums, quantifies, fertility, predicts, penalties, morphologically, tax, iclr, formal, stochastic, map, establishes, conditions, foundations, decouples, output, vocabularies, log, relationship, independently, superword, downstream, superbpe, patching, revives, blt, operating, characters, canine, byt5, wordpiece, units, pair, encoding, sennrich, morfessor, two, dominant, schemes, infersent, cove, doc2vec, sense2vec, handle, oov, static, problem, each, subsection, linking, mace, checks, assisted, 300, example, potato, modal, studio, collecting, curating, rich, assertion, unlimited, documents, foss
Text of the page (random words):
lective state space models linear time long context alternative to attention rwkv rnn transformer hybrid scaling to large parameter counts jamba 2024 hybrid mamba transformer moe architecture rope and yarn rotary position embeddings and context length extension position interpolation extending context windows with minimal fine tuning lost in the middle long context degradation patterns in nlp tasks rag vs long context llms 2024 tradeoffs for qa over long inputs titans learning to memorize at test time 2025 neural long term memory module that learns to memorize historical context at test time scales beyond 2m tokens outperforms transformers and modern linear recurrent models on language modeling and reasoning minimax 01 2025 456b parameter hybrid combining lightning linear attention with sparse softmax attention matches gpt 4o level nlp performance at up to 4m token inference contexts native sparse attention nsa 2025 trainable sparse attention combining coarse grained compression with fine grained selection large speedups at 64k with no nlp benchmark degradation longrope2 2025 identifies undertraining of high frequency rope dimensions and applies evolutionary search rescaling extends llama3 8b to 128k with 80x fewer training tokens than meta s recipe characterizing ssm and hybrid lm long context performance 2025 first comprehensive memory and speed analysis of transformer ssm and hybrid models up to 220k tokens ssms are up to 4x faster hybrids balance recall and efficiency factuality hallucination calibration survey of hallucination in natural language generation taxonomy and mitigation strategies truthfulqa benchmark for truthfulness in question answering factscore fine grained factual precision in long form generation longfact safe 2024 long form factuality benchmark and search augmented evaluator selfcheckgpt sampling based hallucination detection ragas reference free evaluation for rag and qa pipelines lookback lens 2024 attention pattern based hallucination detection in long context generation calibration of llms on multiple choice 2024 calibration analysis under format effects hallulens 2025 hallucination benchmark with extrinsic intrinsic taxonomy and dynamic test set regeneration to resist data leakage atomic calibration 2025 claim level calibration analysis for long form generation models are substantially worse calibrated on extended outputs than on single claims franq 2025 faithfulness aware uncertainty quantification for rag fact checking formally separates faithfulness from factuality much 2025 multilingual claim hallucination benchmark across english french spanish german with token level logits released for principled uq evaluation halluhard 2026 hard multi turn hallucination benchmark for citation required responses 30 hallucination rates persist even with web search cure think through uncertainty 2026 trains models to reason about claim level uncertainty before generating large gains on biography factuality and factbench auroc probing and interpretability a primer in bertology what bert learns about language probing classifiers belinkov methodology limitations alternatives locating and editing factual associations in gpt rome causal tracing of factual recall the pyramid of nlp probes structural probing for linguistic knowledge toy models of superposition foundation for the sparse feature view of transformer representations towards monosemanticity scaling monosemanticity anthropic 2024 sparse autoencoders extracting interpretable features from production scale lms sparse autoencoders find highly interpretable features sae methodology for lm interpretability neuronpedia open platform browsing sae features across models influence functions scale to llms 2023 identifying training examples driving model behavior circuit tracing revealing computational graphs in language models anthropic 2025 introduces cross layer transcoders and attribution graphs to construct an interpretable replacement model enables prompt level circuit tracing of feature to feature causal interactions on the biology of a large language model anthropic 2025 applies attribution graphs to claude 3 5 haiku across multi hop reasoning rhyme planning and jailbreak case studies transcoders beat sparse autoencoders for interpretability 2025 shows transcoders reconstructing layer outputs from inputs yield more interpretable features than saes introduces skip transcoders survey on sparse autoencoders for llm interpretability emnlp 2025 reference survey of sae architectures training strategies feature explanation and evaluation finding highly interpretable prompt specific circuits 2026 identifies circuits at the per prompt level rather than per task reveals mechanism clustering by prompt family efficient and small language models distillation and small models distilbert and minilm distilled encoders for production nlp phi 3 phi 4 microsoft 2024 small models trained on curated data competitive with much larger ones on nlp benchmarks smollm2 huggingface 2025 fully open small lm family with reproducible training data smollm3 huggingface 2025 3b fully open decoder pretrained on 11 2t tokens with nope and yarn for 128k context competitive with 4b class models gemma 3 technical report google 2025 1b 27b open models with high local to global attention ratio to keep kv cache tractable at 128k context qwen3 technical report alibaba 2025 dense and moe models 0 6b 235b with unified thinking non thinking modes the 30b a3b moe matches larger dense models while activating only 3b parameters apple intelligence foundation language models apple 2025 on device 3b model using kv cache sharing and 2 bit qat for 37 5 cache memory reduction without accuracy loss sentence transformers sentence and paragraph embeddings via siamese bert setfit few shot text classification without prompts fastfit fast few shot classification for many class settings gte bge and stella compact text embedding models near the top of mteb quantization and serving relevant when deploying nlp models at scale gptq post training quantization for transformers awq activation aware weight quantization kvtuner icml 2025 sensitivity aware layer wise mixed precision kv cache quantization up to 21 throughput improvement over uniform kv8 gguf llama cpp portable quantized inference vllm pagedattention based high throughput lm serving sglang structured generation and efficient serving text generation inference tgi hf production serving for lms parameter efficient fine tuning lora and qlora low rank adapters and quantized fine tuning the standard for adapting lms to nlp tasks on modest hardware dora 2024 weight decomposed low rank adaptation peft huggingface library bundling lora prefix tuning ia3 and others instruction tuning and preference optimization flan finetuned language models as zero shot learners instructgpt training lms to follow instructions with human feedback self instruct bootstrapping instruction data from lms super naturalinstructions 1600 nlp tasks with instructions constitutional ai training lms with ai generated feedback against a written constitution direct preference optimization simpler alternative to rlhf widely adopted tülu 3 ai2 2024 fully open post training recipe with state of the art results among open models lima less is more for alignment small high quality sft data goes a long way trl reference library for sft dpo grpo and rlhf magpie 2024 2025 synthesizes high quality instruction response pairs by prompting aligned lms with nothing sft on the filtered subset matches official llama 3 instruct bias fairness safety in nlp stereoset measuring stereotypical bias in pretrained lms crows pairs social bias measurement in masked lms winobias gender bias in coreference resolution holisticbias bias measurement across many demographic axes realtoxicityprompts toxicity in lm generation sycophancy in language models models tailoring answers to user beliefs alignment faking in large language models anthropic 2024 models strategically complying during training wildguard 2024 open safety moderation model and benchmark emergent misalignment 2025 finetuning on a narrow task insecure code unexpectedly produces broad alignment failures across unrelated domains safedialbench 2025 multilingual chinese english safety benchmark of 4000 multi turn dialogues across 22 scenarios and 7 jailbreak strategies teleai safety 2025 modular jailbreak evaluation framework integrating 19 attacks 29 defenses and 19 evaluation methods across 14 models and 12 risk categories indicsafe 2026 multilingual safety benchmark across 12 indic languages reveals 12 8 cross language agreement with over refusal in low resource scripts vlaf value conflict alignment faking 2026 alignment faking occurs in models as small as 7b in 37 of cases when policy conflicts with internalized values steering vector mitigation reduces it 94 nlp per language back to top resources organized by human language click a section to expand nlp in arabic back to top libraries camel tools python toolkit for arabic nlp including dialect id morphology ner goarabic go package for arabic text processing jsastem javascript arabic stemmer pyarabic python library for arabic rftokenizer trainable segmenter for arabic hebrew and coptic farasa qcri segmentation pos tagging and ner for arabic models and embeddings arabert arabic bert family camelbert bert models for msa dialectal and classical arabic araelectra efficient arabic pretraining released alongside arabert jais 2023 2024 bilingual arabic english open lm family allam sdaia 2024 arabic first foundation models datasets multidomain datasets largest available multi domain arabic sentiment analysis resources labr large arabic book reviews dataset arabic stopwords aggregated arabic stopwords arabicmmlu 2024 arabic mmlu benchmark nlp in chinese back to top libraries jieba python package for chinese word segmentation snownlp python package for chinese nlp fudannlp java library for chinese text processing hanlp multilingual nlp library with strong chinese support ltp hit language technology platform segmentation pos ner parsing models and embeddings chinese bert wwm whole word masking bert for chinese macbert improved chinese bert with mlm as correction pretraining qwen 2 5 qwen 3 alibaba s open chinese strong lm family chatglm3 glm 4 tsinghua s bilingual chinese english lms baichuan 2 open chinese lm yi 01 ai s bilingual open lms deepseek v3 efficient open moe model with strong chinese anthology funnlp large collection of chinese nlp tools and resources nlp in danish back to top named entity recognition for danish danlp nlp resources in danish awesome danish curated list of resources for danish language technology nlp in dutch back to top python frog python binding to frog an nlp suite for dutch pos tagging lemmatization dependency parsing ner simplenlg_nl dutch surface realiser for natural language generation based on the simplenlg implementation alpino dependency parser for dutch also does pos tagging and lemmatization kaldi nl dutch speech recognition models based on kaldi spacy dutch model industrial strength nlp with a dutch pipeline nlp in german back to top german nlp curated list of open access open source and off the shelf resources and tools developed with a focus on german nlp in hungarian back to top awesome hungarian nlp curated list of free resources for hungarian nlp nlp in indic languages back to top data corpora and treebanks hindi dependency treebank a multi representational multi layered treebank for hindi and urdu universal dependencies treebank in hindi parallel universal dependencies treebank in hindi a smaller part of the above mentioned treebank isi fire stopwords list hindi and bangla peter graham s stopwords list nltk corpus 60k words pos tagged bangla hindi marathi telugu hindi movie reviews dataset 1k samples 3 polarity classes bbc news hindi dataset 4 3k samples 14 classes iit patna hindi absa dataset 5 4k samples 12 domains 4k aspect terms aspect and sentence level polarity in 4 classes bangla absa 5 5k samples 2 domains 10 aspect terms iit patna movie review sentiment dataset 2k samples 3 polarity labels corpora datasets that need a login access can be gained via email sail 2015 twitter and facebook labelled sentiment samples in hindi bengali tamil telugu iit bombay cfilt resources sentiwordnet parallel labelled corpora sense annotated corpora and marathi polarity labelled corpus tdil ic aggregates a lot of useful resources and provides access to otherwise gated datasets language models and word embeddings hindi2vec and nlp for hindi ulmfit style languge model iit patna bilingual word embeddings hi en fasttext word embeddings in a whole bunch of languages trained on common crawl hindi and bengali word2vec hindi and urdu elmo model sanskrit albert trained on sanskrit wikipedia and oscar corpus libraries and tooling multi task deep morphological analyzer deep morphological parser for hindi and urdu indic nlp library tokenization transliteration mt helpers across 18 indic languages sivareddy s dependency parser python3 port dependency parsing and pos tagging for kannada hindi and telugu inltk nlp toolkit for indic languages on pytorch fastai ai4bharat indicnlp suite tools datasets and models across 22 indic languages models and embeddings indicbert v2 2022 2024 multilingual bert for 23 indic languages indictrans2 2023 2024 high quality mt for 22 indic languages openhathi sarvam ai 2023 bilingual hindi english llama continuation airavata 2024 instruction tuned hindi llm sarvam 1 2024 multilingual lm trained from scratch on 10 indic languages bharatgpt krutrim 2024 indic focused foundation models nlp in indonesian back to top libraries and embeddings bahasa natural language toolkit for indonesian indonesian word embedding indonesian fasttext trained on wikipedia pysastrawi python stemmer for bahasa indonesia based on the sastrawi stemming algorithm models indobert indonlu pretrained indonesian lm with the indonlu benchmark suite indobert indolem alternative indobert with the indolem benchmark nusacrowd cendol 2023 2024 large scale community datasets and cendol instruction tuned lms for indonesian and regional languages sailor open southeast asian lms covering indonesian sea lion 2024 singapore ai s open southeast asian lm with strong indonesian datasets kompas and tempo collections at ilps panl10n for pos tagging 39k sentences and 900k word tokens idn for pos tagging 10k sentences and 250k word tokens indonesian treebank and universal dependencies indonesian indosum text summarization and classification wordnet bahasa large free semantic dictionary seacrowd a multilingual and multimodal data hub providing standardized datasets and benchmarks for southeast asian nlp emnlp 2024 nlp in korean back to top libraries konlpy python package for korean natural language processing mecab korean c library for korean nlp koalanlp scala library for korean nlp konlp r package for korean nlp kss ko...
|