Meta tags:
Headings (most frequently used words):
neural, processing, unit, contents, history, use, programming, notes, see, also, references, external, links, consumer, devices, datacenters, scientific, computation,
Text of the page (most frequently used words):
the (47), and (33), for (19), #neural (16), from (16), retrieved (16), july (15), unit (14), edit (14), #processing (13), 2026 (13), use (12), url (11), are (11), cs1 (10), status (10), 2017 (10), npu (10), used (10), maint (9), hardware (9), accelerator (9), link (9), cite (9), web (9), wikipedia (8), 2024 (8), computing (8), gpu (8), nvidia (8), intel (8), such (8), this (7), with (7), computer (7), processors (7), archived (7), original (7), operations (7), using (6), page (6), specific (6), vision (6), matrix (6), apple (6), 2016 (6), google (6), gpus (6), consumer (6), can (6), they (6), fp64 (6), may (5), integrated (5), processor (5), data (5), machine (5), acceleration (5), 2025 (5), performance (5), multiplication (5), november (5), amd (5), also (5), part (5), precision (5), int8 (5), emulated (5), devices (5), models (5), contents (4), search (4), terms (4), articles (4), learning (4), related (4), applications (4), links (4), pdf (4), tensor (4), new (4), chips (4), huawei (4), 2022 (4), silicon (4), movidius (4), march (4), 2012 (4), circuit (4), npus (4), higher (4), low (4), fp16 (4), scientific (4), accelerators (4), history (4), hide (4), move (4), sidebar (4), toggle (3), mobile (3), view (3), september (3), all (3), potentially (3), october (3), english (3), deep (3), units (3), application (3), design (3), memory (3), network (3), chip (3), system (3), level (3), technology (3), platform (3), khronos (3), doi (3), 2023 (3), h100 (3), june (3), architecture (3), qualcomm (3), tops (3), august (3), lake (3), power (3), general (3), purpose (3), approximate (3), programs (3), graphics (3), see (3), cpu (3), large (3), fast (3), apis (3), interfaces (3), file (3), specialized (3), built (3), have (3), trained (3), networks (3), programming (3), designed (3), inference (3), artificial (3), tools (3), main (3), languages (2), table (2), contact (2), about (2), privacy (2), policy (2), organization (2), wikimedia (2), commons (2), was (2), categories (2), containing (2), dated (2), statements (2), 2021 (2), american (2), january (2), 2019 (2), short (2), description (2), wikidata (2), systems (2), logic (2), synthesis (2), emulation (2), digital (2), manycore (2), dataflow (2), architectures (2), asic (2), high (2), custom (2), random (2), number (2), gpgpu (2), software (2), next (2), metal (2), external (2), global (2), current (2), standardization (2), group (2), core (2), ozaki (2), arxiv (2), international (2), faster (2), than (2), amazon (2), how (2), acm (2), datacenter (2), guide (2), vpu (2), gen (2), meteor (2), future (2), here (2), into (2), drones (2), adrian (2), february (2), reality (2), reveals (2), center (2), transistors (2), speed (2), unveils (2), compute (2), references (2), parts (2), not (2), due (2), notes (2), separate (2), field (2), many (2), reduce (2), types (2), opencl (2), vulkan (2), lower (2), tpu (2), cuda (2), vendor (2), coreml (2), which (2), library (2), although (2), multiplications (2), much (2), focus (2), native (2), has (2), since (2), update (2), feature (2), version (2), computation (2), often (2), include (2), dedicated (2), cloud (2), board (2), top (2), datacenters (2), small (2), when (2), per (2), second (2), efficient (2), more (2), circa (2), accelerating (2), that (2), algorithms (2), resolution (2), fps (2), tasks (2), intelligence (2), appearance (2), upload (2), changes (2), read (2), article (2), log (2), create (2), account (2), donate (2), menu (2), add, topic, cookie, statement, statistics, developers, code, conduct, legal, safety, contacts, disclaimers, text, available, under, additional, apply, site, you, agree, registered, trademark, non, profit, foundation, inc, creative, attribution, sharealike, license, rendered, parsoid, last, edited, utc, hidden, unfit, mdy, dates, written, different, gate, arrays, optimization, coprocessors, circuits, https, org, index, php, title, neural_processing_unit, oldid, 1376059740, embedded, virtualization, electronics, chronology, programmable, neuromorphic, systolic, array, heterogeneous, multicore, transport, triggered, cpld, fpga, hdl, implementations, networking, scrypt, attack, tls, cryptography, generation, signal, audio, directx, distributed, parallel, universal, turing, theory, massachusetts, institute, eyeriss, project, puts, pascal, tavenrath, markus, ict, standards, conference, brower, cole, gunnels, john, lopez, graham, technical, blog, unlocking, floating, point, cublas, ootomo, hiroyuki, katsuhisa, yokota, rio, dgemm, integer, 313, 1177, 10943420241239588, 2306, 11975, 297, journal, patel, dylan, nishball, daniel, xie, myron, semianalysis, china, circumvent, restrictions, h20, ascend, 910b, science, innovation, became, secret, sauce, behind, aws, success, jouppi, norman, 1145, 3140659, 3080246, 1704, 04760, sigarch, news, analysis, burns, peter, april, metrics, kan, michael, pcmag, bring, 14th, xdna, lunar, arriving, iphone, connor, fred, 2015, pcworld, microsoft, dives, deeper, hololens, details, holographic, role, revealed, dji, brings, two, flagship, lineup, featuring, myriad, vpus, weckler, independent, dublin, tech, firm, virtual, headset, snapdragon, ifa, moss, sebastian, dynamics, hopper, billion, merritt, rick, times, designing, geekom, ultra, choose, comprehensive, esmaeilzadeh, hadi, sampson, ceze, luis, burger, doug, 460, 1109, micro, 449, 45th, annual, ieee, symposium, microarchitecture, wiggers, kyle, 2020, venturebeat, magic, raises, million, boost, inferencing, off, shelf, insidehpc, inspur, gx4, carly, stick, usb, customized, task, type, ray, tracing, composed, organic, material, wetware, mlx, builds, atop, engine, ane, relatively, good, unified, there, underlying, compilers, runtimes, causing, great, increase, development, effort, combinations, involved, open, standard, pursuing, amount, work, needed, working, three, fronts, expansion, intrinsic, inclusion, graphs, skriptnd, format, describing, nnef, spir, generally, existing, pipelines, adapted, precisions, being, private, accessible, through, snpe, each, their, own, upon, openvino, ryzen, operating, provide, litert, android, ios, macos, windows, formats, represent, onnx, directml, tensorflow, tailored, emulate, modern, place, making, scheme, outperform, been, demonstrated, a100, especially, benefited, limited, capacity, showing, speedup, toolkit, automatically, uses, equivalent, addition, fp32, introduced, blas, titan, rtx, late, 2010s, companies, form, functional, these, commonly, both, training, servers, tpus, exist, category, without, dominant, emerging, services, inferentia, trainium, package, plus, hbm, stacks, printed, pcb, liquid, cooled, packages, front, panel, side, pcie, connectors, intended, but, reasonably, run, support, bitwidth, int4, common, metric, trillions, does, explicitly, specify, kind, typically, additions, fp8, recently, added, versatile, cnn, sift, need, keep, track, objects, visually, headsets, augmented, scale, invariant, transform, convolutional, smartphone, rendering, significantly, resource, allowing, traditional, render, scene, frame, rate, then, pre, model, turn, base, imagery, smoother, output, 2160p, 240, real, time, frames, 540p, samsung, intensive, sensor, driven, designs, arithmetic, novel, capability, widely, grade, mosfets, oxide, semiconductor, effect, contains, tens, billions, spatial, internet, things, robotics, either, efficiently, execute, already, like, llms, train, consumption, language, early, term, refer, appeared, paper, described, known, class, accelerate, including, standalone, central, module, attached, via, adapter, hat, raspberry, hailo, redirected, free, encyclopedia, item, other, projects, printable, download, print, export, switch, legacy, parser, get, shortened, information, permanent, what, actions, talk, tiếng, việt, українська, shqip, русский, română, português, polski, 한국어, 日本語, italiano, bahasa, indonesia, français, suomi, فارسی, eesti, español, deutsch, čeština, català, беларуская, العربية, subsection, personal, special, pages, recent, community, portal, learn, help, contribute, events, navigation, jump, content,
Text of the page (random words):
neural processing unit wikipedia jump to content main menu main menu move to sidebar hide navigation main page contents current events random article about wikipedia contact us contribute help learn to edit community portal recent changes upload file special pages search search appearance donate create account log in personal tools donate create account log in contents move to sidebar hide top 1 history 2 use toggle use subsection 2 1 consumer devices 2 2 datacenters 2 3 scientific computation 3 programming 4 notes 5 see also 6 references 7 external links toggle the table of contents neural processing unit 22 languages العربية беларуская català čeština deutsch español eesti فارسی suomi français bahasa indonesia italiano 日本語 한국어 polski português română русский shqip українська tiếng việt 中文 edit links article talk english read edit view history tools tools move to sidebar hide actions read edit view history general what links here related changes upload file permanent link page information cite this page get shortened url switch to legacy parser print export download as pdf printable version in other projects wikimedia commons wikidata item appearance move to sidebar hide from wikipedia the free encyclopedia redirected from ai accelerator hardware acceleration unit for artificial intelligence tasks a hailo ai accelerator module attached to a raspberry pi 5 via an m 2 adapter hat 2024 a neural processing unit npu also known as an ai accelerator or deep learning processor is a class of specialized hardware accelerator 1 or computer system 2 3 designed to accelerate artificial intelligence and machine learning applications including artificial neural networks and computer vision npu can be standalone a part of a central processing unit cpu or a part of a graphics processing unit gpu history edit an early use of the term neural processing unit npu to refer to a dedicated neural network accelerator appeared in the 2012 paper neural acceleration for general purpose approximate programs which described an npu architecture for accelerating approximate programs 4 use edit npu s purpose is either to efficiently execute already trained ai models like large language models llms for inference or to train ai models npus can be more efficient in terms of speed or power consumption 5 npu applications include algorithms for robotics internet of things and data intensive or sensor driven tasks 6 they are often manycore or spatial designs and focus on low precision arithmetic novel dataflow architectures or in memory computing capability as of 2024 update a widely used datacenter grade ai integrated circuit chip the nvidia h100 gpu contains tens of billions of metal oxide semiconductor field effect transistors mosfets 7 consumer devices edit ai accelerators are used in apple silicon qualcomm samsung huawei 8 and google tensor smartphone processors 9 when used as part of a gpu for graphics rendering they can significantly reduce resource use by allowing the traditional parts of the gpu to render a scene at a much lower resolution and frame rate e g 540p at 30 frames per second fps and then using a pre trained ai model on the npu to turn that base imagery into smoother and higher resolution output e g 2160p at 240 fps in real time vision processing units are accelerators specialized for machine vision algorithms such as convolutional neural networks cnn and scale invariant feature transform sift they are used in devices that need to keep track of objects visually such as augmented reality ar headsets and drones 10 11 12 it is more recently circa 2017 added to processors from apple 13 and circa 2022 to processors from intel 14 and amd 15 all models of intel meteor lake processors have a built in versatile processor unit vpu for accelerating inference for computer vision and deep learning 16 on consumer devices the npu is intended to be small power efficient but reasonably fast when used to run small models to do this they are designed to support low bitwidth operations using data types such as int4 int8 fp8 and fp16 a common metric is trillions of operations per second tops although tops does not explicitly specify the kind of operations it is typically int8 additions and multiplications 17 datacenters edit the google tensor processing unit tpu v4 package asic in center plus 4 hbm stacks and printed circuit board pcb with 4 liquid cooled packages the board s front panel has 4 top side pcie connectors 2023 accelerators are used in cloud computing servers e g tpus for google cloud platform 18 and trainium and inferentia chips for amazon web services 19 many vendor specific terms exist for devices in this category and it is an emerging technology without a dominant design since the late 2010s gpus designed by companies such as nvidia and amd often include ai specific hardware in the form of dedicated functional units for low precision matrix multiplication operations these gpus are commonly used as ai accelerators both for training and inference 20 scientific computation edit see also scientific computing although npus are tailored for low precision e g fp16 int8 matrix multiplication operations they can be used to emulate higher precision matrix multiplications in scientific computing as modern gpus place much focus on making the npu part fast using emulated fp64 ozaki scheme on npus can potentially outperform native fp64 this has been demonstrated using fp16 emulated fp64 on nvidia titan rtx and using int8 emulated fp64 on nvidia consumer gpus and the a100 gpu consumer gpus especially benefited as they have limited fp64 hardware capacity showing a 6 speedup 21 since cuda toolkit 13 0 update 2 cu blas automatically uses int8 emulated fp64 matrix multiplication of the equivalent precision if it is faster than native this is in addition to the fp16 emulated fp32 feature introduced in version 12 9 22 programming edit an operating system or a higher level library may provide application programming interfaces such as tensorflow with litert next android coreml ios macos or directml windows formats such as onnx are used to represent trained neural networks consumer cpu integrated npus are accessible through vendor specific apis amd ryzen ai intel openvino apple silicon coreml a and qualcomm snpe each have their own apis which can be built upon by a higher level library gpus generally use existing gpgpu pipelines such as cuda and opencl adapted for lower precisions and specialized matrix multiplication operations vulkan is also being used custom built systems such as the google tpu use private interfaces there are a large number of separate underlying acceleration apis and compilers runtimes in use in the ai field causing a great increase in software development effort due to the many combinations involved as of 2025 the open standard organization khronos group is pursuing standardization of ai related interfaces to reduce the amount of work needed khronos is working on three separate fronts expansion of data types and intrinsic operations in opencl and vulkan inclusion of compute graphs in spir v and a nnef skriptnd file format for describing a neural network 23 notes edit mlx builds atop the cpu and gpu parts not the apple neural engine ane part of apple silicon chips the relatively good performance is due to the use of a large fast unified memory design see also edit wetware computer computer composed of organic material ray tracing hardware type of 3d graphics accelerator application specific integrated circuit integrated circuit customized for a specific task references edit page carly july 21 2017 intel unveils movidius compute stick usb ai accelerator v3 archived from the original on august 11 2017 retrieved august 11 2017 inspur unveils gx4 ai accelerator insidehpc june 21 2017 archived from the original on june 21 2017 wiggers kyle november 6 2019 neural magic raises 15 million to boost ai inferencing speed on off the shelf processors venturebeat archived from the original on march 6 2020 esmaeilzadeh hadi sampson adrian ceze luis burger doug 2012 neural acceleration for general purpose approximate programs 2012 45th annual ieee acm international symposium on microarchitecture pp 449 460 doi 10 1109 micro 2012 48 intel core vs ultra how to choose comprehensive guide geekom march 24 2025 retrieved july 5 2026 merritt rick may 18 2016 google designing ai processors ee times retrieved july 5 2026 cite web cs1 maint url status link moss sebastian march 23 2022 nvidia reveals new hopper h100 gpu with 80 billion transistors data center dynamics retrieved july 5 2026 cite web cs1 maint url status link huawei reveals the future of mobile ai at ifa 2017 huawei global september 2 2017 archived from the original on november 10 2021 retrieved january 28 2024 snapdragon 8 gen 3 mobile platform pdf archived from the original pdf on october 25 2023 weckler adrian february 14 2016 dublin tech firm movidius to power google s new virtual reality headset independent ie archived from the original on february 19 2016 retrieved march 15 2016 dji brings two new flagship drones to lineup featuring myriad 2 vpus machine vision technology movidius movidius november 15 2016 archived from the original on november 26 2016 o connor fred may 1 2015 microsoft dives deeper into hololens details holographic processor role revealed pcworld retrieved july 5 2026 cite web cs1 maint url status link the future is here iphone x apple september 12 2017 retrieved july 5 2026 intel s lunar lake processors arriving q3 2024 intel may 20 2024 retrieved july 5 2026 cite web cs1 maint url status link amd xdna architecture amd retrieved july 5 2026 cite web cs1 maint url status link kan michael august 1 2022 intel to bring a vpu processor unit to 14th gen meteor lake chips pcmag retrieved july 5 2026 cite web cs1 maint url status link burns peter april 24 2024 a guide to ai tops and npu performance metrics qualcomm retrieved july 5 2026 cite web cs1 maint url status link jouppi norman p et al june 24 2017 in datacenter performance analysis of a tensor processing unit acm sigarch computer architecture news 45 2 1 12 arxiv 1704 04760 doi 10 1145 3140659 3080246 how silicon innovation became the secret sauce behind aws s success amazon science july 27 2022 retrieved july 5 2026 patel dylan nishball daniel xie myron november 9 2023 nvidia s new china ai chips circumvent us restrictions h20 faster than h100 huawei ascend 910b semianalysis retrieved july 5 2026 ootomo hiroyuki ozaki katsuhisa yokota rio july 2024 dgemm on integer matrix multiplication unit the international journal of high performance computing applications 38 4 297 313 arxiv 2306 11975 doi 10 1177 10943420241239588 brower cole gunnels john lopez graham october 24 2025 unlocking tensor core performance with floating point emulation in cublas nvidia technical blog retrieved july 5 2026 cite web cs1 maint url status link tavenrath markus 2025 current status of ai related standardization in the khronos group pdf global ict standards conference 2025 external links edit nvidia puts the accelerator to the metal with pascal the next platform eyeriss project massachusetts institute of technology v t e hardware acceleration theory universal turing machine parallel computing distributed computing applications gpu gpgpu software directx audio digital signal processing hardware random number generation neural processing unit cryptography tls machine vision custom hardware attack scrypt networking data implementations high level synthesis c to hdl fpga asic cpld system on a chip network on a chip architectures dataflow transport triggered multicore manycore heterogeneous in memory computing systolic array neuromorphic related programmable logic processor design chronology digital electronics virtualization hardware emulation logic synthesis embedded systems retrieved from https en wikipedia org w index php title neural_processing_unit oldid 1376059740 categories application specific integrated circuits neural processing units coprocessors computer optimization gate arrays deep learning hidden categories articles with short description short description is different from wikidata use american english from january 2019 all wikipedia articles written in american english use mdy dates from october 2021 cs1 unfit url articles containing potentially dated statements from 2024 all articles containing potentially dated statements cs1 maint url status this page was last edited on 21 september 2026 at 20 09 utc page was rendered with parsoid text is available under the creative commons attribution sharealike 4 0 license additional terms may apply by using this site you agree to the terms of use and privacy policy wikipedia is a registered trademark of the wikimedia foundation inc a non profit organization privacy policy about wikipedia disclaimers contact wikipedia legal safety contacts code of conduct developers statistics cookie statement mobile view search search toggle the table of contents neural processing unit 22 languages add topic
|