If you are not sure if the website you would like to visit is secure, you can verify it here. Enter the website address of the page and see parts of its content and the thumbnail images on this site. None (if any) dangerous scripts on the referenced page will be executed. Additionally, if the selected site contains subpages, you can verify it (review) in batches containing 5 pages.
favicon.ico: en.wikipedia.org/wiki/Post-training_of_large_language_models - Post-training of large languag.

site address: en.wikipedia.org/wiki/Post-training_of_large_language_models redirected to: en.wikipedia.org/wiki/Post-training_of_large_language_models

site title: Post-training of large language models - Wikipedia

Our opinion (on Monday 05 October 2026 13:29:06 UTC):

GREEN status (no comments) - no comments
After content analysis of this website we propose the following hashtags:



Meta tags:

Headings (most frequently used words):

and, training, tuning, preference, learning, feedback, post, of, large, language, models, contents, scope, terminology, development, methods, data, evaluation, limitations, see, also, references, supervised, fine, instruction, alignment, reasoning, focused, parameter, efficient, adaptation, distillation, reward, modeling, reinforcement, from, human, direct, optimization, ai, generated,

Text of the page (most frequently used words):
and (62), the (58), training (42), model (42), learning (41), from (27), language (26), feedback (25), post (22), #models (22), preference (21), data (19), reasoning (19), human (19), arxiv (19), tuning (19), can (19), for (18), instruction (18), edit (18), reward (17), with (16), reinforcement (15), large (14), neural (14), optimization (13), machine (13), systems (13), policy (12), used (12), methods (11), network (10), supervised (10), distillation (10), generated (10), fine (10), that (10), alignment (9), wikipedia (8), use (8), research (8), task (8), stage (8), behavior (8), tasks (8), comparisons (8), may (7), processing (7), based (7), self (7), 2024 (7), international (7), conference (7), information (7), proceedings (7), also (7), base (7), objective (7), its (7), outputs (7), response (7), datasets (6), advances (6), low (6), adaptation (6), direct (6), instructions (6), following (6), are (6), judgments (6), trained (6), page (5), knowledge (5), artificial (5), llm (5), parameter (5), 2023 (5), efficient (5), lora (5), rank (5), step (5), depend (5), evaluation (5), recent (5), surveys (5), rather (5), than (5), labels (5), while (5), move (5), sft (5), add (4), contents (4), search (4), text (4), this (4), was (4), generative (4), john (4), general (4), open (4), rlhf (4), gradient (4), method (4), representations (4), quantized (4), 2022 (4), scaling (4), limitations (4), multi (4), stages (4), learned (4), quality (4), some (4), uses (4), responses (4), process (4), outcomes (4), examples (4), such (4), results (4), not (4), focused (4), several (4), hide (4), sidebar (4), toggle (3), view (3), safety (3), about (3), under (3), using (3), all (3), written (3), english (3), short (3), wikidata (3), natural (3), applications (3), long (3), term (3), memory (3), inference (3), diffusion (3), image (3), deep (3), automated (3), loss (3), history (3), qlora (3), 2305 (3), rlaif (3), vol (3), computational (3), kto (3), generalization (3), 2025 (3), into (3), increase (3), comparison (3), performance (3), these (3), uniformly (3), correctness (3), labeler (3), single (3), prompts (3), matched (3), intermediate (3), steps (3), development (3), separate (3), between (3), differ (3), target (3), checked (3), coverage (3), but (3), trains (3), include (3), train (3), larger (3), parameters (3), rationales (3), one (3), studied (3), establish (3), desirable (3), undesirable (3), dpo (3), common (3), pipeline (3), account (3), after (3), pretraining (3), scope (3), tools (3), main (3), languages (2), table (2), contact (2), privacy (2), available (2), additional (2), terms (2), apply (2), attribution (2), last (2), july (2), 2026 (2), categories (2), articles (2), american (2), description (2), different (2), impact (2), visual (2), video (2), effect (2), center (2), act (2), autoencoder (2), rnn (2), recurrent (2), vision (2), transformer (2), computer (2), turing (2), architectures (2), alex (2), fei (2), stephen (2), paul (2), people (2), ibm (2), recognition (2), speech (2), synthesis (2), protocol (2), intelligence (2), laws (2), agent (2), context (2), approach (2), symbolic (2), source (2), engineering (2), regularization (2), weight (2), descent (2), constraint (2), projects (2), algorithms (2), minillm (2), twelfth (2), eric (2), star (2), 2203 (2), 235 (2), 41st (2), rafailov (2), rafael (2), instruct (2), association (2), linguistics (2), transactions (2), ouyang (2), jeff (2), 2020 (2), multitask (2), prompted (2), quantization (2), survey (2), references (2), see (2), pipelines (2), cost (2), difficult (2), design (2), sampling (2), choices (2), recipes (2), among (2), diversity (2), every (2), signal (2), evaluator (2), update (2), settings (2), each (2), held (2), out (2), evaluated (2), answer (2), during (2), unseen (2), reduce (2), supervision (2), filtering (2), invalid (2), student (2), selected (2), teacher (2), divergence (2), smaller (2), adapters (2), frozen (2), through (2), reducing (2), describe (2), updates (2), fixed (2), trainable (2), retained (2), preferable (2), within (2), level (2), automatically (2), preferences (2), produced (2), experiments (2), summarization (2), related (2), objectives (2), example (2), labeled (2), without (2), across (2), preferred (2), form (2), classification (2), depends (2), reference (2), predict (2), collected (2), resulting (2), modeling (2), candidate (2), then (2), commonly (2), later (2), converted (2), which (2), mixtures (2), scale (2), pairs (2), work (2), optimized (2), lines (2), time (2), label (2), terminology (2), appearance (2), upload (2), file (2), changes (2), links (2), read (2), article (2), log (2), create (2), donate (2), menu (2), topic, mobile, cookie, statement, statistics, developers, code, conduct, legal, contacts, disclaimers, site, you, agree, registered, trademark, non, profit, organization, wikimedia, foundation, inc, creative, commons, sharealike, license, rendered, parsoid, edited, utc, hidden, retrieved, https, org, index, php, title, training_of_large_language_models, oldid, 1363846749, category, workplace, warfare, military, art, games, marketing, chatbot, psychosis, healthcare, fiction, education, architecture, engine, explainable, environmental, competition, arms, race, anthropomorphism, winter, slop, literacy, infrastructure, bubble, boom, social, economic, opposition, centers, propaganda, virtual, politician, regulation, precautionary, principle, nationalism, ethics, elections, takeover, government, cold, war, political, graph, gnn, adversarial, gan, variational, vae, mamba, highway, residual, convolutional, cnn, multilayer, perceptron, mlp, echo, state, gated, unit, gru, lstm, vit, differentiable, françois, chollet, daniel, kokotajlo, jan, leike, mustafa, suleyman, schulman, aidan, gomez, noam, shazeer, ashish, vaswani, andrej, karpathy, david, silver, demis, hassabis, ian, goodfellow, quoc, oriol, vinyals, ilya, sutskever, krizhevsky, andrew, james, goodnight, graves, grossberg, lotfi, zadeh, yoshua, bengio, yann, lecun, jürgen, schmidhuber, hopfield, geoffrey, hinton, werbos, seppo, linnainmaa, seymour, papert, joseph, weizenbaum, bernard, widrow, frank, rosenblatt, oliver, selfridge, herbert, simon, cliff, shaw, allen, newell, nathaniel, rochester, mccarthy, marvin, minsky, takeo, kanade, kunihiko, fukushima, shun, ichi, amari, claude, shannon, christopher, manning, von, neumann, walter, pitts, warren, sturgis, mcculloch, alan, yago, dbpedia, conceptnet, bases, opencog, lida, clarion, soar, cognitive, rule, semantic, reasoners, procedural, logic, programs, engines, expert, deductive, classifiers, robot, control, autogpt, action, selection, muzero, driving, car, openai, five, alphazero, alphago, decisional, watsonx, watson, project, debater, list, oasis, genie, world, udio, suno, riffusion, music, generation, veo, seedance, sora, kling, hailuo, runway, gen, dream, stable, recraft, midjourney, imagen, ideogram, gpt, flux, firefly, dall, aurora, alphafold, facial, whisper, elevenlabs, ocr, hwr, wavenet, alexnet, audio, implementations, physical, agent2agent, hypothetical, superintelligence, asi, agi, weak, lethal, autonomous, weapons, humanity, exam, companion, intelligent, nmt, game, playing, theorem, proving, actor, critic, algorithm, situated, neuro, sovereign, blended, vibe, coding, word, embedding, hallucination, recursive, improvement, reflection, uncanny, valley, rag, adversary, autoregression, latent, imitation, sarsa, prompt, augmentation, initialization, gating, rectifier, sigmoid, softmax, activation, batchnorm, normalization, convolution, attention, backpropagation, conjugate, quasi, newton, sgd, clustering, overfitting, double, bias, variance, tradeoff, regression, functions, hyperparameter, representation, satisfaction, planning, concepts, lists, software, proprietary, institutions, companies, glossary, timeline, yuxian, dong, wei, furu, huang, minlie, 2306, 08543, dettmers, tim, pagnoni, artidoro, holtzman, ari, zettlemoyer, luke, finetuning, llms, 14314, edward, shen, yelong, wallis, phillip, 2106, 09685, zelikman, yuhuai, jesse, goodman, noah, bootstrapping, 14465, lee, harrison, phatale, samrat, mansoor, hassan, 26901, 2309, 00267, 26874, chittepu, yaswanth, park, ryan, overoptimization, 2406, 02900, wang, yizhong, kordi, yeganeh, mishra, swaroop, aligning, 2212, 10560, 61st, annual, meeting, ethayarajh, kawin, winnie, muennighoff, niklas, prospect, theoretic, 2402, 01306, sharma, archit, mitchell, your, secretly, 18290, casper, davies, xander, shi, claudia, problems, fundamental, 2307, 15217, jiang, follow, 02155, stiennon, nisan, summarize, 2009, 01325, sanh, victor, webson, albert, raffel, colin, enables, zero, shot, 2110, 08207, nagel, markus, amjad, rana, ali, van, baalen, mart, 119, 7206, 7197, 37th, down, adaptive, rounding, lambert, nathan, morrison, jacob, pyatkin, valentina, tulu, pushing, frontiers, 2411, 15124, lightman, hunter, kosaraju, vineet, burda, yura, let, verify, 20050, kaufmann, timo, weng, bengs, viktor, hüllermeier, eyke, 2312, 14925, chung, hyung, won, hou, longpre, shayne, finetuned, 2210, 11416, journal, kumar, komal, ashraf, tajamul, thawakar, omkar, dive, 2502, 21321, tie, guiyao, zhao, zeli, song, dingjie, 2503, 06072, make, causal, final, curation, order, limited, disclosure, hinder, independent, reproduction, sequential, interfere, previously, discuss, catastrophic, forgetting, capability, regressions, trade, offs, concerns, ordering, imperfect, inconsistent, unrepresentative, obtain, whose, hard, judge, overfit, implicit, measured, continue, improve, external, plateaus, deteriorates, score, measures, goals, improvements, fit, domain, differently, same, sensitive, decoding, intended, assessed, compared, where, benchmarks, evaluations, afterward, benchmark, decontamination, overlap, desirability, rules, tests, matching, synthetic, expand, remove, repetitive, generations, reproduce, transferred, token, probabilities, reverse, kullback, leibler, teachers, deployment, consolidate, remains, dependent, combines, backpropagates, gradients, bit, how, parameterized, they, constituting, behavioral, full, require, substantial, accelerator, set, added, leaving, most, weights, represents, matrices, number, storage, needed, adaptations, taught, reasoner, iteratively, led, correct, answers, tuned, instead, evaluates, study, mathematical, outperformed, outcome, tested, setting, specific, recipe, worked, solutions, problem, solving, assisted, another, dialogue, comparable, baselines, reduces, dependence, capabilities, criteria, guide, assumptions, designed, alternative, alternatives, evidence, learns, directly, rejected, derivation, expresses, optimal, class, regularized, fitted, binary, still, strength, begins, collects, those, optimizes, predicted, proxy, measurement, universal, annotators, aggregation, construct, only, rankings, pairwise, probability, subject, limits, often, appears, beginning, formats, basic, before, provide, generates, verification, existing, synthetically, assembled, sources, generate, inputs, filtered, near, duplicate, individual, amount, composition, works, vary, distribution, procedure, given, input, organized, paired, desired, examines, construction, introduced, reformulated, avoided, separately, online, kahneman, tversky, subsequently, proposed, requiring, developed, alongside, abstractive, against, instructgpt, demonstrations, ranked, reviews, combination, collection, now, described, emerged, collections, excluded, mixture, flan, families, sizes, chain, thought, increased, computation, broad, others, sense, distinct, lowers, numerical, precision, already, teaching, umbrella, including, oriented, efficiency, their, boundaries, does, imply, universally, agreed, sequence, centered, usage, pretrained, further, structured, practical, combine, followed, technical, literature, applied, initial, varies, elements, broader, treatments, whether, tool, integration, included, free, encyclopedia, item, other, printable, version, download, pdf, print, export, switch, legacy, parser, get, shortened, url, cite, permanent, link, what, here, actions, talk, subsection, top, personal, special, pages, community, portal, learn, help, contribute, random, current, events, navigation, jump, content,


Text of the page (random words):
earch appearance donate create account log in personal tools donate create account log in contents move to sidebar hide top 1 scope and terminology 2 development 3 methods toggle methods subsection 3 1 supervised fine tuning and instruction tuning 3 2 preference learning and alignment 3 2 1 reward modeling and reinforcement learning from human feedback 3 2 2 direct preference optimization 3 2 3 ai generated feedback 3 3 reasoning focused training 3 4 parameter efficient adaptation 3 5 distillation 4 data and evaluation 5 limitations 6 see also 7 references toggle the table of contents post training of large language models add languages add links article talk english read edit view history tools tools move to sidebar hide actions read edit view history general what links here related changes upload file permanent link page information cite this page get shortened url switch to legacy parser print export download as pdf printable version in other projects wikidata item appearance move to sidebar hide from wikipedia the free encyclopedia training methods used after llm pretraining post training of large language models is a term used in recent technical literature for training applied to a large language model llm after its initial large scale pretraining the scope varies among surveys supervised fine tuning or instruction tuning and preference based alignment are common elements while broader treatments differ on whether reasoning focused training parameter efficient adaptation distillation tool integration or inference time scaling are included 1 2 in training centered usage a pretrained or base model is further optimized with structured data such as instruction response pairs comparisons between candidate responses step level judgments or outcomes that can be checked automatically practical systems may combine several stages for example supervised training followed by preference optimization or reinforcement learning 3 4 5 6 scope and terminology edit recent surveys use post training as an umbrella label for several lines of research including instruction tuning preference alignment reasoning oriented training efficiency methods and distillation their boundaries differ so the label does not imply a single objective or a universally agreed sequence of stages 1 2 some surveys include increased computation at inference time in a broad account of post training while others center the term on additional training after pretraining 1 2 post training in this sense is also distinct from post training quantization which lowers the numerical precision of an already trained network rather than teaching instruction following preferences or reasoning behavior 7 development edit the methods now described as post training emerged through several lines of work 1 2 multitask instruction tuning trained models on collections of tasks written as natural language prompts t0 converted supervised datasets into prompted tasks and evaluated generalization to tasks excluded from its training mixture 8 later flan experiments studied instruction tuning across larger task mixtures model families model sizes and chain of thought data 3 preference based training developed alongside instruction tuning work on abstractive summarization trained a reward model from human comparisons and optimized a language model against that reward 9 instructgpt used supervised fine tuning on demonstrations reward model training on ranked outputs and reinforcement learning from human feedback 10 reviews of reinforcement learning from human feedback describe the approach as a combination of feedback collection reward learning and policy optimization with limitations at each stage 4 11 direct preference optimization dpo introduced in 2023 reformulated a commonly used preference learning objective as a classification loss and avoided the separately trained reward model and online policy gradient stage used in a common reinforcement learning from human feedback pipeline 12 kahneman tversky optimization kto was subsequently proposed for data labeled desirable or undesirable without requiring matched response pairs 13 methods edit supervised fine tuning and instruction tuning edit in supervised fine tuning sft the model is trained to predict a target response for a given input instruction tuning is a form of sft in which examples are organized as natural language instructions paired with desired outputs research on instruction tuning examines task mixtures model scale response construction data quality and generalization to unseen tasks 3 instruction data can be written by people converted from existing task datasets generated synthetically or assembled from several sources self instruct used a language model to generate instructions inputs and outputs filtered invalid or near duplicate examples and then used the resulting data for fine tuning 14 results from individual recipes do not establish a fixed amount or composition of data that works for every model or task outcomes vary with the base model task distribution and training procedure 3 sft often appears at the beginning of a multi stage pipeline it can establish response formats and basic instruction following before preference optimization or reinforcement learning and it can provide a policy that generates outputs for later comparison or verification 10 6 preference learning and alignment edit preference learning uses judgments about candidate outputs rather than only a single target response feedback may be collected as rankings pairwise comparisons or labels such as desirable and undesirable the model is then trained to increase the probability of preferred behavior commonly subject to a constraint that limits divergence from a reference policy 4 10 12 reward modeling and reinforcement learning from human feedback edit a common reinforcement learning from human feedback rlhf pipeline begins with sft collects comparisons between model outputs trains a reward model to predict those comparisons and optimizes the language model policy to increase the predicted reward 10 9 4 the reward model is a learned proxy for the collected feedback rather than a direct measurement of a universal human preference the resulting behavior depends on the annotators instructions sampling process and aggregation method used to construct the feedback data 10 11 direct preference optimization edit dpo learns directly from preferred and rejected responses its derivation expresses the optimal policy for a class of reward regularized objectives in a form that can be fitted with a binary classification loss 12 dpo still depends on the coverage and quality of comparison data and on choices such as the reference policy and regularization strength 12 15 related objectives use different feedback assumptions kto for example was designed for examples labeled desirable or undesirable without a matched alternative response 13 these methods are alternatives within preference based training not evidence that one objective is uniformly preferable across datasets and applications 13 11 ai generated feedback edit feedback can also be generated or assisted by another language model reinforcement learning from ai feedback rlaif trains a reward model using preferences produced by an ai labeler in experiments on summarization and dialogue tasks rlaif produced results comparable to the studied rlhf baselines 16 ai generated feedback reduces dependence on human labels but its results depend on the capabilities and prompts of the labeler model and on the criteria used to guide its judgments 16 11 reasoning focused training edit some recent surveys include reasoning focused methods within post training 2 1 these methods use worked solutions generated rationales step level feedback or automatically checked outcomes to train multi step problem solving the self taught reasoner star iteratively generated rationales retained rationales that led to correct answers and fine tuned on the retained examples 17 process supervision instead evaluates intermediate reasoning steps in one study of mathematical reasoning process supervised reward models outperformed outcome supervised reward models in the tested setting 5 such results are specific to the studied tasks and models and do not establish that one reasoning training recipe is uniformly preferable 2 parameter efficient adaptation edit full fine tuning updates all model parameters and can require substantial accelerator memory parameter efficient methods train a smaller set of added or selected parameters while leaving most base model weights fixed low rank adaptation lora represents weight updates with trainable low rank matrices reducing the number of trainable parameters and the storage needed for separate task adaptations 18 quantized low rank adaptation qlora combines low rank adapters with a quantized frozen base model its design backpropagates gradients through a frozen 4 bit quantized model into lora adapters reducing memory use during fine tuning 19 lora and qlora describe how an update is parameterized they can be used with instruction tuning or preference optimization rather than constituting a separate behavioral objective 18 19 distillation edit knowledge distillation trains a student model to reproduce selected behavior from a teacher model in llm post training the transferred signal may include token probabilities generated responses or instruction following behavior 1 minillm used an on policy distillation objective based on reverse kullback leibler divergence to train smaller generative language models from larger teachers 20 distillation can reduce deployment cost or consolidate behavior learned by a larger model but the student remains dependent on the teacher outputs and on the coverage and filtering of the distillation data 20 1 data and evaluation edit post training datasets differ by stage sft uses target responses preference optimization uses comparisons or desirability labels process supervision uses judgments on intermediate steps and some reinforcement learning settings use outcomes that can be checked by rules tests or answer matching 3 4 5 synthetic data can expand coverage but filtering is used to remove invalid repetitive or low quality generations 14 evaluation is matched to the intended effect of each stage instruction following can be assessed on held out tasks and human judgments preference models can be compared on held out comparisons reasoning systems can be evaluated using answer correctness and where available judgments on intermediate steps 3 4 5 multi stage pipelines may use development benchmarks during training and separate unseen evaluations afterward some also apply benchmark decontamination to reduce overlap between training and evaluation data 6 no single score measures all post training goals improvements in instruction following preference fit reasoning correctness diversity safety or domain performance can move differently under the same update comparisons are also sensitive to the base model training data prompts decoding settings and evaluator 4 11 15 limitations edit preference based methods depend on imperfect feedback human labels can be inconsistent unrepresentative or difficult to obtain for tasks whose correctness is hard to judge while ai generated labels depend on the labeler model and its instructions 4 11 16 optimization can also overfit a learned or implicit preference signal performance measured by the training objective may continue to improve while quality under an external evaluator plateaus or deteriorates 15 sequential stages can interfere with previously learned behavior recent post training surveys discuss catastrophic forgetting capability regressions and trade offs among alignment diversity and general task performance 1 2 these concerns depend on the base model data objective and ordering of stages rather than following uniformly from every post training method multi stage pipelines increase computational and engineering cost and make causal attribution difficult final behavior may depend on the base model data curation stage order reward design sampling and evaluation choices 6 4 limited disclosure of post training datasets and recipes can hinder independent reproduction and comparison 6 see also edit ai alignment fine tuning deep learning knowledge distillation large language model lora machine learning reasoning model reinforcement learning from human feedback references edit 1 2 3 4 5 6 7 8 tie guiyao zhao zeli song dingjie et al 2025 a survey on post training of large language models arxiv 2503 06072 cs cl 1 2 3 4 5 6 7 kumar komal ashraf tajamul thawakar omkar et al 2025 llm post training a deep dive into reasoning large language models arxiv 2502 21321 cs cl 1 2 3 4 5 6 chung hyung won hou le longpre shayne et al 2024 scaling instruction finetuned language models journal of machine learning research 25 70 1 53 arxiv 2210 11416 1 2 3 4 5 6 7 8 9 kaufmann timo weng paul bengs viktor hüllermeier eyke 2025 a survey of reinforcement learning from human feedback transactions on machine learning research arxiv 2312 14925 1 2 3 4 lightman hunter kosaraju vineet burda yura et al 2024 let s verify step by step the twelfth international conference on learning representations arxiv 2305 20050 1 2 3 4 5 lambert nathan morrison jacob pyatkin valentina et al 2024 tulu 3 pushing frontiers in open language model post training arxiv 2411 15124 cs cl nagel markus amjad rana ali van baalen mart et al 2020 up or down adaptive rounding for post training quantization proceedings of the 37th international conference on machine learning proceedings of machine learning research vol 119 pp 7197 7206 sanh victor webson albert raffel colin et al 2022 multitask prompted training enables zero shot task generalization international conference on learning representations arxiv 2110 08207 1 2 stiennon nisan ouyang long wu jeff et al 2020 learning to summarize from human feedback advances in neural information processing systems 33 arxiv 2009 01325 1 2 3 4 5 ouyang long wu jeff jiang xu et al 2022 training language models to follow instructions with human feedback advances in neural information processing systems 35 arxiv 2203 02155 1 2 3 4 5 6 casper stephen davies xander shi claudia et al 2023 open problems and fundamental limitations of reinforcement learning from human feedback transactions on machine learning research arxiv 2307 15217 1 2 3 4 rafailov rafael sharma archit mitchell eric et al 2023 direct preference optimization your language model is secretly a reward model advances in neural information processing systems 36 arxiv 2305 18290 1 2 3 ethayarajh kawin xu winnie muennighoff niklas et al 2024 kto model alignment as prospect theoretic optimization proceedings of the 41st international conference on machine learning proceedings of machine learning research vol 235 arxiv 2402 01306 1 2 wang yizhong kordi yeganeh mishra swaroop et al 2023 self instruct alig...
Thumbnail images (randomly selected): * Images may be subject to copyright.GREEN status (no comments)
  • Wikipedia
  • The Free Encyclopedia
  • Wikimedia Foundation
  • Powered by MediaWiki

Verified site has: 321 subpage(s). Do you want to verify them? Verify pages:

1-5 6-10 11-15 16-20 21-25 26-30 31-35 36-40 41-45 46-50
51-55 56-60 61-65 66-70 71-75 76-80 81-85 86-90 91-95 96-100
101-105 106-110 111-115 116-120 121-125 126-130 131-135 136-140 141-145 146-150
151-155 156-160 161-165 166-170 171-175 176-180 181-185 186-190 191-195 196-200
201-205 206-210 211-215 216-220 221-225 226-230 231-235 236-240 241-245 246-250
251-255 256-260 261-265 266-270 271-275 276-280 281-285 286-290 291-295 296-300
301-305 306-310 311-315 316-320 321-321


The site also has references to the 1 subdomain(s)

  en.wikipedia.org  Verify


The site also has 2 references to other resources (not html/xhtml )

 en.wikipedia.org/wiki/15.ai  Verify  stats.wikimedia.org/#/en.wikipedia.org  Verify


Top 50 hastags from of all verified websites.

Supplementary Information (add-on for SEO geeks)*- See more on header.verify-www.com

Header

HTTP/1.1 301 Moved Permanently
content-length 0
location htt????/en.wikipedia.org/wiki/Post-training_of_large_language_models
server HAProxy
x-cache cp6011 int
x-cache-status int-tls
connection close
HTTP/2 200
date Mon, 05 Oct 2026 13:29:06 GMT
server mw-web.eqiad.main-6dbc6d9d5-lgnvx
x-content-type-options nosniff
content-language en
accept-ch
reporting-endpoints csp-report-to-endpoint= /w/api.php?action=cspreport&format=json ;
content-security-policy script-src unsafe-eval blob: self meta.wikimedia.org *.wikimedia.org *.wikipedia.org *.wikinews.org *.wiktionary.org *.wikibooks.org *.wikiversity.org *.wikisource.org wikisource.org *.wikiquote.org *.wikidata.org *.wikifunctions.org *.wikivoyage.org *.mediawiki.org mediawiki.org wikimedia.org *.wmflabs.org *.wmcloud.org *.toolforge.org wss://*.toolforge.org *.jsdelivr.net unpkg.com cdnjs.cloudflare.com raw.githubusercontent.com *.github.com code.jquery.com cdn.mathjax.org use.typekit.net fonts.cdnfonts.com use.fontawesome.com i.ytimg.com rsms.me doi.org localhost htt????/localhost:* htt???/localhost:* wss://localhost:* ws://localhost:* *.google.com *.gstatic.com *.googleapis.com *.translate.yandex.net yastatic.net ya.ru radically.github.io cdn.sammdot.ca cdn.fontshare.com viaf.org publicai-proxy.alaexis.workers.dev iiif.archive.org api.flickr.com live.staticflickr.com api.anthropic.com api.openai.com api.publicai.co catalogo.pusc.it parsifal.urbe.it opac.sbn.it overpass-api.de api.openrouteservice.org archive.org *.openstreetmap.org *.waymarkedtrails.org *.thunderforest.com registry.ipe.wiki analytics.ipe.wiki qlever.dev app.goacoustic.com wikipedia-archive.ourworldindata.org api.inaturalist.org inaturalist-open-data.s3.amazonaws.com validator.w3.org db.onlinewebfonts.com fontlibrary.org unsafe-inline auth.wikimedia.org; default-src self data: blob: upload.wikimedia.org thumb.wikimedia.org htt????/commons.wikimedia.org meta.wikimedia.org *.wikimedia.org *.wikipedia.org *.wikinews.org *.wiktionary.org *.wikibooks.org *.wikiversity.org *.wikisource.org wikisource.org *.wikiquote.org *.wikidata.org *.wikifunctions.org *.wikivoyage.org *.mediawiki.org mediawiki.org wikimedia.org *.wmflabs.org *.wmcloud.org *.toolforge.org wss://*.toolforge.org *.jsdelivr.net unpkg.com cdnjs.cloudflare.com raw.githubusercontent.com *.github.com code.jquery.com cdn.mathjax.org use.typekit.net fonts.cdnfonts.com use.fontawesome.com i.ytimg.com rsms.me doi.org localhost htt????/localhost:* htt???/localhost:* wss://localhost:* ws://localhost:* *.google.com *.gstatic.com *.googleapis.com *.translate.yandex.net yastatic.net ya.ru radically.github.io cdn.sammdot.ca cdn.fontshare.com viaf.org publicai-proxy.alaexis.workers.dev iiif.archive.org api.flickr.com live.staticflickr.com api.anthropic.com api.openai.com api.publicai.co catalogo.pusc.it parsifal.urbe.it opac.sbn.it overpass-api.de api.openrouteservice.org archive.org *.openstreetmap.org *.waymarkedtrails.org *.thunderforest.com registry.ipe.wiki analytics.ipe.wiki qlever.dev app.goacoustic.com wikipedia-archive.ourworldindata.org api.inaturalist.org inaturalist-open-data.s3.amazonaws.com validator.w3.org db.onlinewebfonts.com fontlibrary.org en.wikibooks.org en.wikinews.org en.wikiquote.org en.wikisource.org en.wikiversity.org en.wikivoyage.org en.wiktionary.org www.mediawiki.org commons.wikimedia.org foundation.wikimedia.org incubator.wikimedia.org species.wikimedia.org wikimania.wikimedia.org www.wikidata.org www.wikifunctions.org auth.wikimedia.org; style-src self data: blob: upload.wikimedia.org thumb.wikimedia.org htt????/commons.wikimedia.org meta.wikimedia.org *.wikimedia.org *.wikipedia.org *.wikinews.org *.wiktionary.org *.wikibooks.org *.wikiversity.org *.wikisource.org wikisource.org *.wikiquote.org *.wikidata.org *.wikifunctions.org *.wikivoyage.org *.mediawiki.org mediawiki.org wikimedia.org *.wmflabs.org *.wmcloud.org *.toolforge.org wss://*.toolforge.org *.jsdelivr.net unpkg.com cdnjs.cloudflare.com raw.githubusercontent.com *.github.com code.jquery.com cdn.mathjax.org use.typekit.net fonts.cdnfonts.com use.fontawesome.com i.ytimg.com rsms.me doi.org localhost htt????/localhost:* htt???/localhost:* wss://localhost:* ws://localhost:* *.google.com *.gstatic.com *.googleapis.com *.translate.yandex.net yastatic.net ya.ru radically.github.io cdn.sammdot.ca cdn.fontshare.com viaf.org publicai-proxy.alaexis.workers.dev iiif.archive.org api.flickr.com live.staticflickr.com api.anthropic.com api.openai.com api.publicai.co catalogo.pusc.it parsifal.urbe.it opac.sbn.it overpass-api.de api.openrouteservice.org archive.org *.openstreetmap.org *.waymarkedtrails.org *.thunderforest.com registry.ipe.wiki analytics.ipe.wiki qlever.dev app.goacoustic.com wikipedia-archive.ourworldindata.org api.inaturalist.org inaturalist-open-data.s3.amazonaws.com validator.w3.org db.onlinewebfonts.com fontlibrary.org unsafe-inline ; object-src none ; report-uri /w/api.php?action=cspreport&format=json; report-to csp-report-to-endpoint
last-modified Tue, 29 Sep 2026 19:54:36 GMT
content-type text/html; charset=UTF-8
content-encoding gzip
age 1
accept-ranges bytes
x-cache cp6015 miss, cp6009 miss
x-cache-status miss
strict-transport-security max-age=106384710; includeSubDomains; preload
report-to group : wm_nel , max_age : 604800, endpoints : [ url : htt????/intake-logging.wikimedia.org/v1/events?stream=w3c.reportingapi.network_error&schema_uri=/w3c/reportingapi/network_error/1.0.0 ]
nel report_to : wm_nel , max_age : 604800, failure_fraction : 0.05, success_fraction : 0.0
set-cookie WMF-Last-Access=05-Oct-2026;Path=/;HttpOnly;secure;Expires=Fri, 06 Nov 2026 12:00:00 GMT
set-cookie WMF-Last-Access-Global=05-Oct-2026;Path=/;Domain=.wikipedia.org;HttpOnly;secure;Expires=Fri, 06 Nov 2026 12:00:00 GMT
set-cookie WMF-DP=36b;Path=/;HttpOnly;secure;Expires=Tue, 06 Oct 2026 00:00:00 GMT
x-client-ip 5.135.42.194
cache-control private, s-maxage=0, max-age=0, must-revalidate, no-transform
vary Accept-Encoding,X-Subdomain,Cookie,Authorization,User-Agent
set-cookie GeoIP=FR:::48.86:2.34:v4; Path=/; secure; Domain=.wikipedia.org
set-cookie NetworkProbeLimit=0.001;Path=/;Secure;SameSite=None;Max-Age=3600
set-cookie WMF-Uniq=corsfqMpgK_q7V8AZUMJgAPwAAAAAFvdTXI0rmAU328ryRJ9Q5GUdrVCY4H9lUNI;Domain=.wikipedia.org;Path=/;HttpOnly;secure;SameSite=None;Expires=Tue, 05 Oct 2027 00:00:00 GMT
x-request-id 3eb7e022-3a53-48aa-9675-d54cd48f56f4
x-analytics
server-timing cache;desc= miss , host;desc= cp6009 ,co_id;desc= 2744354313

Meta Tags

title="Post-training of large language models - Wikipedia"
charset="UTF-8"
name="ResourceLoaderDynamicStyles" content=""
name="generator" content="MediaWiki 1.47.0-wmf.22"
name="referrer" content="origin"
name="referrer" content="origin-when-cross-origin"
name="robots" content="noindex,nofollow,max-image-preview:standard"
name="format-detection" content="telephone=no"
name="viewport" content="width=1120"
property="og:title" content="Post-training of large language models - Wikipedia"
property="og:type" content="website"
property="mw:PageProp/toc" id="mwIw" data-mw='{"autoGenerated":true}'

Load Info

page size218130
load time (s)0.52277
redirect count1
speed download72128
server IP 185.15.58.224
* all occurrences of the string "http://" have been changed to "htt???/"