Meta tags:
Headings (most frequently used words):
cache, in, policies, associative, multi, first, microprocessors, cpu, history, operation, two, way, example, and, caches, write, instruction, versus, contents, overview, associativity, entry, structure, miss, address, translation, hierarchy, modern, processor, implementation, see, also, notes, references, external, links, entries, performance, direct, mapped, set, speculative, execution, skewed, pseudo, multicolumn, flag, bits, homonym, synonym, problems, virtual, tags, hints, page, coloring, specialized, level, scratchpad, memory, the, k8, more, hierarchies, tag, ram, ported, replacement, stalls, victim, trace, coalescing, wcc, micro, μop, or, uop, branch, target, smart, core, chips, separate, unified, exclusive, inclusive, tlb, implementations, data, 68k, x86, arm, current, research,
Text of the page (most frequently used words):
the (933), cache (604), and (288), #memory (176), for (145), data (125), #caches (110), with (107), that (103), from (100), address (90), virtual (89), are (84), which (79), instruction (77), this (70), have (67), can (66), cpu (62), associative (62), processor (60), main (60), physical (59), way (58), edit (58), one (58), used (57), level (55), has (54), set (53), not (49), intel (48), tag (45), core (44), performance (43), two (42), each (41), bits (39), bit (38), instructions (37), access (36), retrieved (35), tlb (35), also (35), read (35), system (34), index (33), pdf (32), kib (32), may (31), was (31), chip (31), write (31), there (31), some (31), page (30), entry (30), but (30), into (28), time (28), location (28), more (28), than (28), use (27), computer (27), size (27), processors (27), only (27), tags (27), other (26), when (26), ibm (25), mapped (25), first (25), miss (25), line (25), all (24), different (24), unit (24), multi (24), larger (24), same (23), pages (23), block (23), high (22), direct (22), power (21), architecture (21), been (21), original (21), speed (21), addresses (21), mib (21), both (21), blocks (21), hierarchy (20), archived (20), number (20), per (19), between (19), latency (19), because (19), however (19), hardware (18), its (18), associativity (18), micro (18), most (18), any (18), lines (18), hit (18), cores (18), cpus (17), branch (17), translation (17), shared (17), they (17), multiple (17), design (16), register (16), store (16), doi (16), operation (16), amd (16), were (16), much (16), virtually (16), then (16), locations (16), program (15), systems (15), article (15), sram (15), does (15), fetch (15), indexed (15), entries (15), misses (15), policy (14), processing (14), four (14), possible (14), space (14), physically (14), such (14), usually (14), generally (14), called (14), replacement (13), policies (13), single (13), execution (13), machine (13), another (13), part (13), 2013 (13), operating (13), had (13), levels (13), early (13), these (13), faster (13), byte (13), since (13), tagged (13), must (13), will (13), mapping (13), buffer (12), pipeline (12), stored (12), based (12), trace (12), pentium (12), large (12), example (12), least (12), before (12), those (12), table (11), value (11), rate (11), 128 (11), microarchitecture (11), caching (11), very (11), inclusive (11), well (11), fast (11), see (11), where (11), area (11), microprocessors (11), die (11), motherboard (11), split (11), need (11), separate (11), cached (11), μop (11), available (10), microprocessor (10), management (10), mmu (10), low (10), order (10), 360 (10), arm (10), conflict (10), 2012 (10), 2014 (10), small (10), new (10), major (10), would (10), ram (10), through (10), sizes (10), victim (10), after (10), eight (10), modern (10), having (10), thus (10), implemented (10), decoded (10), using (9), non (9), model (9), history (9), 256 (9), several (9), times (9), international (9), 1990 (9), exclusive (9), information (9), case (9), designs (9), advantage (9), slower (9), either (9), above (9), typically (9), ways (9), keep (9), copies (9), copy (9), smaller (9), require (9), many (9), toggle (8), computing (8), control (8), target (8), bus (8), scratchpad (8), cycle (8), secondary (8), 2009 (8), symposium (8), software (8), work (8), cost (8), skewed (8), stores (8), fetched (8), edram (8), common (8), their (8), hints (8), hint (8), although (8), simple (8), like (8), unified (8), accesses (8), rather (8), specialized (8), include (8), coloring (8), problem (8), multicolumn (8), contents (7), wikipedia (7), cs1 (7), 2010 (7), 2008 (7), deprecated (7), service (7), general (7), dynamic (7), application (7), package (7), types (7), second (7), x86 (7), external (7), 1109 (7), proceedings (7), conference (7), 2020 (7), 1145 (7), problems (7), isbn (7), 2015 (7), cite (7), anandtech (7), 2001 (7), length (7), september (7), uses (7), given (7), allows (7), important (7), just (7), dram (7), three (7), bytes (7), while (7), path (7), few (7), function (7), evicted (7), fewer (7), known (7), become (7), color (7), colors (7), tagging (7), requires (7), offset (7), valid (7), flag (7), code (6), last (6), december (6), 2023 (6), org (6), module (6), clock (6), integrated (6), decoder (6), registers (6), predictor (6), load (6), coherence (6), graphics (6), operations (6), cycles (6), prediction (6), random (6), written (6), november (6), ieee (6), s2cid (6), smart (6), haswell (6), fully (6), gap (6), tlbs (6), help (6), total (6), placement (6), effectively (6), recently (6), still (6), significant (6), form (6), extra (6), taken (6), programs (6), needed (6), later (6), determine (6), requested (6), often (6), until (6), whether (6), among (6), map (6), simpler (6), structure (6), run (6), longer (6), hold (6), makes (6), delay (6), particular (6), selected (6), lower (6), below (6), back (6), hash (6), dirty (6), log (6), due (6), subsection (6), search (5), links (5), articles (5), march (5), maint (5), archival (5), central (5), quantum (5), file (5), logic (5), lookaside (5), variable (5), word (5), thread (5), wide (5), issue (5), out (5), risc (5), alpha (5), mips (5), sets (5), turing (5), what (5), 2003 (5), family (5), solutions (5), parallel (5), com (5), introduction (5), link (5), pro (5), 978 (5), bridge (5), atlas (5), over (5), technology (5), reduced (5), llc (5), fetching (5), instead (5), added (5), iii (5), board (5), next (5), implementation (5), being (5), rates (5), could (5), improve (5), right (5), released (5), following (5), provided (5), done (5), various (5), associated (5), checking (5), checked (5), make (5), takes (5), means (5), result (5), better (5), ptes (5), storing (5), allow (5), match (5), get (5), patterns (5), full (5), needs (5), reduce (5), aliases (5), existing (5), vipt (5), contains (5), pseudo (5), evict (5), about (4), additional (4), terms (4), august (4), statements (4), short (4), security (4), electronics (4), purpose (4), signal (4), functional (4), magnitude (4), pipelined (4), speculative (4), process (4), scalar (4), isa (4), itanium (4), sparc (4), series (4), automaton (4), technologies (4), flushing (4), 2004 (4), reference (4), acm (4), xeon (4), 1968 (4), october (4), cortex (4), 2007 (4), hierarchies (4), advanced (4), techniques (4), web (4), iris (4), 2011 (4), 2016 (4), frontend (4), david (4), decode (4), uop (4), optimization (4), zhang (4), slice (4), scheme (4), predicting (4), fact (4), approach (4), works (4), tables (4), every (4), ported (4), accessing (4), benefit (4), tools (4), energy (4), recent (4), current (4), depending (4), type (4), unusually (4), introduced (4), dynamically (4), became (4), included (4), directly (4), chips (4), loop (4), special (4), implementations (4), complex (4), fastest (4), might (4), similar (4), against (4), compared (4), reads (4), diagram (4), indexing (4), reading (4), higher (4), content (4), keeps (4), ensure (4), coherency (4), predictors (4), kinds (4), parity (4), effective (4), maps (4), athlon (4), comparable (4), hits (4), versus (4), highest (4), increasing (4), accessed (4), reducing (4), complexity (4), protocols (4), amount (4), off (4), examples (4), actual (4), without (4), avoiding (4), wcc (4), writes (4), already (4), guarantee (4), alternatively (4), changes (4), mappings (4), difficult (4), within (4), differences (4), best (4), aliasing (4), less (4), hide (4), move (4), sidebar (4), mobile (3), view (3), organization (3), long (3), unsourced (3), array (3), digital (3), related (3), frequency (3), switch (3), sum (3), addressed (3), adder (3), counter (3), sequential (3), controller (3), point (3), generation (3), others (3), 512 (3), network (3), gpu (3), multiprocessor (3), embedded (3), vector (3), updates (3), simultaneous (3), algorithm (3), false (3), sharing (3), stall (3), dec (3), comparison (3), edge (3), specific (3), storage (3), cellular (3), machines (3), queue (3), review (3), updated (3), simulation (3), smith (3), introduces (3), capacity (3), annual (3), product (3), world (3), james (3), overview (3), 486 (3), corporation (3), microcomputer (3), 2nd (3), characteristics (3), joint (3), third (3), february (3), 645 (3), manual (3), technical (3), analysis (3), usa (3), jouppi (3), norman (3), sigarch (3), news (3), tested (3), 2006 (3), april (3), 2018 (3), how (3), real (3), bulldozer (3), programmers (3), compiler (3), buffers (3), 2024 (3), lake (3), multicore (3), computers (3), john (3), selection (3), shown (3), advantages (3), conventional (3), consumption (3), atom (3), 1976 (3), memories (3), intended (3), idea (3), placed (3), loops (3), 2021 (3), simulator (3), algorithms (3), request (3), efficiency (3), apple (3), feature (3), uncommon (3), find (3), took (3), slightly (3), mhz (3), located (3), devices (3), 68k (3), documented (3), especially (3), found (3), smallest (3), largest (3), subset (3), occurred (3), rows (3), pair (3), simplest (3), select (3), sometimes (3), loaded (3), even (3), making (3), addressable (3), coherent (3), ecc (3), predict (3), branches (3), them (3), banks (3), sections (3), ten (3), always (3), local (3), instance (3), intermediate (3), present (3), certain (3), should (3), cray (3), itself (3), down (3), l1i (3), holds (3), providing (3), required (3), traces (3), fixed (3), decreasing (3), coalescing (3), points (3), technique (3), during (3), likely (3), pattern (3), lead (3), necessary (3), described (3), flushed (3), changed (3), becomes (3), solved (3), writing (3), synonym (3), homonym (3), historically (3), lookup (3), choice (3), r6000 (3), looked (3), vivt (3), refer (3), context (3), implement (3), comparing (3), cause (3), continue (3), invalid (3), stale (3), row (3), closer (3), lru (3), transistors (3), places (3), doubling (3), choose (3), waiting (3), stalls (3), checks (3), exception (3), settle (3), topic (2), languages (2), contact (2), privacy (2), wikimedia (2), commons (2), categories (2), wayback (2), disputed (2), accuracy (2), description (2), wikidata (2), pin (2), semiconductor (2), device (2), ppw (2), watt (2), voltage (2), scaling (2), circuit (2), circuitry (2), barrel (2), shifter (2), datapath (2), stack (2), gate (2), floating (2), fifo (2), heterogeneous (2), count (2), slicing (2), image (2), coprocessor (2), stream (2), transistor (2), temporal (2), multithreading (2), task (2), pipelining (2), parallelism (2), renaming (2), structural (2), hazards (2), classic (2), visc (2), motorola (2), addressing (2), architectures (2), finite (2), state (2), models (2), lecture (2), power4 (2), understanding (2), introductory (2), 2002 (2), hill (2), paper (2), results (2), benchmarks (2), net (2), describing (2), cacti (2), foster (2), chen (2), june (2), a22 (2), 6916 (2), 1964 (2), 6600 (2), 1999 (2), bottleneck (2), 123 (2), hotchips (2), cluster (2), tian (2), 5200 (2), 4950hq (2), difference (2), leon (2), cutress (2), ian (2), zen (2), shimpi (2), anand (2), lal (2), islped (2), aware (2), kanter (2), sandy (2), continued (2), agner (2), guide (2), ghz (2), skylake (2), desktop (2), products (2), formerly (2), improving (2), 17th (2), isca (2), jiang (2), lin (2), xiaodong (2), george (2), mechanism (2), journal (2), citeseerx (2), vol (2), 1962 (2), ifip (2), congress (2), patterson (2), hennessy (2), interface (2), choosing (2), error (2), protection (2), patents (2), justia (2), patent (2), issued (2), 1994 (2), reconfigurable (2), 1997 (2), schemes (2), ryan (2), codenamed (2), ivy (2), networks (2), z13 (2), throughput (2), 1982 (2), developed (2), cambridge (2), sequence (2), slave (2), give (2), effect (2), communication (2), references (2), notes (2), allocation (2), locality (2), serve (2), traditional (2), normally (2), whereas (2), connected (2), entirely (2), average (2), consider (2), research (2), 192 (2), laptop (2), variant (2), serves (2), crystalwell (2), again (2), tens (2), popularity (2), socket (2), obsolete (2), began (2), onto (2), sockets (2), made (2), growing (2), versions (2), 386 (2), amounts (2), expensive (2), around (2), left (2), 68040 (2), 68020 (2), mode (2), considered (2), tiny (2), replaced (2), true (2), 1980s (2), introduce (2), price (2), closely (2), semi (2), conductor (2), mainframe (2), 1960s (2), flat (2), extensive (2), values (2), depend (2), greatly (2), needing (2), citation (2), decoders (2), hinted (2), width (2), recurrence (2), translated (2), finally (2), srams (2), shows (2), matching (2), calculated (2), portion (2), retired (2), tends (2), sensitive (2), great (2), silicon (2), currently (2), change (2), guaranteed (2), terminology (2), primary (2), never (2), fill (2), keeping (2), translate (2), marking (2), date (2), specialization (2), here (2), reduces (2), strictly (2), check (2), together (2), disadvantage (2), whenever (2), corresponding (2), quite (2), name (2), separately (2), meaning (2), benefits (2), including (2), ranges (2), independently (2), resulting (2), reasons (2), allowing (2), share (2), question (2), implementing (2), files (2), did (2), end (2), scheduled (2), allocates (2), nest (2), phenom (2), features (2), along (2), returned (2), cases (2), replacing (2), fundamental (2), tradeoff (2), operate (2), dedicated (2), equal (2), method (2), powered (2), delivering (2), enough (2), causing (2), created (2), heuristic (2), l1d (2), basic (2), extreme (2), commonly (2), property (2), role (2), server (2), probes (2), mentioned (2), 6000 (2), attempt (2), assign (2), simply (2), extremely (2), collide (2), arrange (2), literature (2), community (2), cannot (2), consistent (2), timing (2), perhaps (2), suffers (2), determining (2), alternate (2), matches (2), distinguish (2), simultaneously (2), perfect (2), proceed (2), allowed (2), inconsistent (2), overlapping (2), spaces (2), now (2), tend (2), pipt (2), homonyms (2), designed (2), exist (2), potential (2), additionally (2), avoids (2), slow (2), divided (2), held (2), running (2), cut (2), dependent (2), measurement (2), indicates (2), marked (2), implies (2), field (2), hence (2), displaystyle (2), lceil (2), rceil (2), ratio (2), contributes (2), conflicts (2), tests (2), something (2), good (2), suffer (2), hand (2), unpredictable (2), acts (2), free (2), place (2), execute (2), independent (2), waits (2), hundreds (2), able (2), steps (2), immediately (2), rarely (2), future (2), contain (2), copied (2), etc (2), static (2), executable (2), step (2), appearance (2), upload (2), create (2), account (2), donate (2), menu (2), add, cookie, statement, statistics, developers, conduct, legal, safety, contacts, disclaimers, text, under, apply, site, you, agree, registered, trademark, profit, foundation, inc, creative, attribution, sharealike, license, rendered, parsoid, edited, 2026, utc, hidden, webarchive, template, volume, disputes, errors, parameters, https, php, title, cpu_cache, oldid, 1372368516, carrier, grid, tick, tock, fabrication, chronology, gating, acpi, apm, pmu, analog, boolean, mixed, binary, multiplier, demultiplexer, multiplexer, rom, microcode, hardwired, status, glue, combinational, imc, fpu, agu, alu, arithmetic, units, components, manycore, baseband, secure, cryptoprocessor, tpu, tensor, dsp, ppu, physics, vpu, vision, accelerator, accelerators, noc, cypress, psoc, mpsoc, soc, soft, asip, ultra, microcontroller, pop, sip, mcm, cpld, fpoa, fpga, asic, pal, tile, gpgpu, orders, metrics, sups, synaptic, tps, transactions, flops, ips, ipc, cpi, spmd, mimd, misd, swar, simt, simd, sisd, flynn, taxonomy, cooperative, preemptive, heterogenous, hyperthreading, distributed, superscalar, serial, dependence, reservation, station, tomasulo, scoreboarding, dependency, operand, forwarding, epiphany, tilera, 3x0, 390, 370, lmc, microblaze, openrisc, unicore, m32r, etrax, cris, superh, clipper, powerpc, stanford, pdp, vax, 68000, modes, zisc, nisc, oisc, misc, epic, vliw, trips, cisc, orthogonal, neuromorphic, cognitive, multiprocessing, fabric, huma, numa, endianness, transport, triggered, dataflow, modified, harvard, von, neumann, pointer, belt, zeno, hypercomputation, probabilistic, nondeterministic, post, universal, alternating, deterministic, hierarchical, abstract, fallacy, princeton, university, ixbtlabs, pavel, danilov, ars, technica, jon, stokes, vhdl, paul, genua, freescale, primer, ruud, van, der, pas, sun, microsystems, cantin, thorough, lucidly, presented, reasonably, organizations, spec, cpu2000, 1989, compulsory, classification, evaluating, lwn, ulrich, drepper, detail, wikibook, labs, wang, zhenghong, lee, ruby, 41st, novel, enhanced, sally, adee, 43892134, mspec, 5292036, spectrum, thwarts, sneak, attack, ark, specs, july, reilly, kheradpir, shervin, allan, flight, thornton, fall, edition, 1972, ga27, 2719, january, electric, mahapatra, nihar, venkatrao, balakrishna, 3es, 11557476, 357783, 331677, crossroads, l210, r4f, sandpile, jaleel, aamer, eric, borch, bhandaru, malini, steely, simon, emer, joel, jaleels, achieving, zheng, ying, davis, brian, jordan, matthew, austin, texas, 7803, 8385, ispass, 1291359, evaluation, amecomputers, explanation, bradley, borg, anita, 1992, 114, 146628, 139708, study, lempel, oded, cornell, workshop, interconnects, shih, chiu, edaboard, betw, inside, demo, niu, kun, btic, motiani, dipti, 2017, dual, schedulers, revealed, analyzed, solomon, baruch, mendelson, avi, orenstein, doron, almog, yoav, ronen, ronny, huntington, beach, 195859085, 58113, 371, lpe, 945363, association, machinery, subsystem, via, assembly, makers, fog, 2000, launch, crystal, addition, prefetch, seattle, 373, 134547, 364, letter, qingda, ding, xiaoning, zhao, sadayappan, 14th, salt, city, utah, 378, hpca, 4658653, 367, gaining, insights, partitioning, bridging, roscoe, timothy, baumann, andrew, ethz, 263, 3800, 00l, taylor, davies, peter, farmwald, michael, 2si, 363, 325096, 325161, 355, bottomley, linux, parameter, kaxiras, stefanos, ros, alberto, 40th, 547, 15434231, 9781450320795, 2485922, 2485968, 307, 9125, 535, perspective, efficient, kilburn, payne, howarth, 1961, conferences, eastern, washington, macmillan, 294, 279, key, supervisor, sumner, haley, chenh, spartan, neill, proc, afips, spring, 1967, 621, 1465482, 1465581, 611, experience, multiprogramming, relocation, cragon, harvey, 1996, jones, bartlett, learning, 209, 86720, 474, dugan, ben, concerning, cooperman, gene, basics, morgan, kaufmann, 484, 374493, sadler, nathan, sorin, daniel, 20160350229, developer, programmer, chenxi, yan, yong, 621212, 1997imicr, 17e, 40c, bibcode, kozyrakis, seznec, andré, 1993, 178, 173682, 165152, 169, jahagirdar, sanjeev, varghese, sodhi, inder, wells, megalingam, rajesh, kannan, deepu, joseph, iype, vikram, vandana, science, 556, 18236635, 4244, 4519, iccsit, 5234663, 551, phased, ucsd, edu, launches, p5900, 10nm, radio, press, release, 32kb, 5mb, 15mb, newsroom, sheet, accelerating, infrastructure, white, bill, cecilia, z13s, elsevier, 383872, quantitative, altering, raise, suggest, researchers, alan, jay, 530, 6023466, 356887, 356892, 473, surveys, liptay, 1147, 0015, aspects, landy, barry, tunnel, diode, worked, speeded, operands, obeyed, modulo, remaining, wanted, speedup, words, mathematical, laboratory, aldermaston, cad, centre, chao, zeng, qingkai, nicopolitidis, petros, 1939, 0122, issn, 1155, 5559552, survey, side, channel, attacks, systematic, countermeasures, torres, gabriel, paging, frame, ferranti, memoization, dinero, prefetching, ports, phases, concept, super, architects, explore, tradeoffs, simplescalar, open, source, options, focused, fault, tolerance, goals, mainframes, equipped, gt3e, increasingly, newer, generations, megabytes, enjoyed, prolonged, thanks, previously, named, motherboards, produced, disappeared, development, brought, clocked, termed, differentiate, boards, contained, 485turbocache, kbyte, era, disparity, caused, sdram, mmx, daughtercard, support, reached, featured, 120, refresh, constructed, significantly, latencies, nine, enable, optional, upgrade, dip, cells, bottom, corner, a38202, austek, simms, i386, 68060, kilobytes, 1987, basically, shrink, burst, 68030, accelerates, consist, 1984, typical, 68010, cdc, days, technically, economically, viable, plenty, alleviate, combined, tied, invention, scarcity, span, magnetic, drum, disc, seen, ahead, studies, optimize, optimal, programming, language, algol, fortran, cobol, discuss, improved, delays, collapsing, looks, matched, supplies, similarly, adjacent, clarify, manner, fields, multiplexing, complicated, chooses, save, element, relevant, extracted, returns, aligned, bypassed, inner, sure, restarted, deal, effort, expended, engineering, specify, employ, relied, show, calls, facility, costly, compute, discussing, speaks, thought, bypass, 21264, interesting, trick, protected, accidental, corruption, strike, spare, particle, usual, class, fairly, targets, jumps, pte, responsible, portions, spread, handles, identical, section, fetches, boundaries, predecoding, damaged, fresh, illustrate, spm, internal, temporary, calculations, progress, swapped, saved, incremental, wish, remove, enforce, inclusion, drawback, correlation, associativities, restricted, eviction, possibly, maintain, inclusiveness, diminishes, hitting, exchanged, exchange, copying, decisions, somewhere, reside, universally, accepted, names, partially, demonstrated, constraint, referred, pieces, undesirable, increase, considerably, global, desirable, whole, redundancy, processes, threads, utilized, considering, inevitably, wiring, circa, usable, characteristic, assignments, reallocated, runtime, bank, break, dependencies, easing, acting, essentially, tulsa, incorporated, compatible, madison, 1995, 21164, incorporating, begun, utilize, pull, entire, 2010s, mounted, fourth, rare, 2019, flip, flop, starting, oryon, qualcomm, lunar, a14, z15
Text of the page (random words):
example loop nest optimization however with register renaming most compiler register assignments are reallocated dynamically by hardware at runtime into a register bank allowing the cpu to break false data dependencies and thus easing pipeline hazards register files sometimes also have hierarchy the cray 1 circa 1976 had eight address a and eight scalar data s registers that were generally usable there was also a set of 64 address b and 64 scalar data t registers that took longer to access but were faster than main memory the b and t registers were provided because the cray 1 did not have a data cache the cray 1 did however have an instruction cache multi core chips edit when considering a chip with multiple cores there is a question of whether the caches should be shared or local to each core implementing shared cache inevitably introduces more wiring and complexity but then having one cache per chip rather than core greatly reduces the amount of space needed and thus one can include a larger cache typically sharing the l1 cache is undesirable because the resulting increase in latency would make each core run considerably slower than a single core chip however for the highest level cache usually l3 the last one called before accessing memory having a global cache is desirable for several reasons such as allowing a single core to use the whole cache reducing data redundancy by making it possible for different processes or threads to share cached data and reducing the complexity of utilized cache coherency protocols 54 for example an eight core chip with three levels may include an l1 cache for each core one intermediate l2 cache for each pair of cores and one l3 cache shared between all cores a shared highest level cache usually l3 called before accessing memory is usually referred to as a last level cache llc 55 additional techniques are used for increasing the level of parallelism when llc is shared between multiple cores including slicing it into multiple pieces which are addressing certain ranges of memory addresses and can be accessed independently 8 56 separate versus unified edit in a separate cache structure instructions and data are cached separately meaning that a cache line is used to cache either instructions or data but not both various benefits have been demonstrated with separate data and instruction translation lookaside buffers 57 in a unified structure this constraint is not present and cache lines can be used to cache both instructions and data exclusive versus inclusive edit multi level caches introduce new design decisions for instance in some processors all data in the l1 cache must also be somewhere in the l2 cache these caches are called strictly inclusive other processors like the amd athlon have exclusive caches data are guaranteed to be in at most one of the l1 and l2 caches never in both still other processors like the intel pentium ii iii and 4 do not require that data in the l1 cache also reside in the l2 cache although it may often do so there is no universally accepted name for this intermediate policy 58 59 two common names are non exclusive and partially inclusive the advantage of exclusive caches is that they store more data this advantage is larger when the exclusive l1 cache is comparable to the l2 cache and diminishes if the l2 cache is many times larger than the l1 cache when the l1 misses and the l2 hits on an access the hitting cache line in the l2 is exchanged with a line in the l1 this exchange is quite a bit more work than just copying a line from l2 to l1 which is what an inclusive cache does 59 one advantage of strictly inclusive caches is that when external devices or other processors in a multiprocessor system wish to remove a cache line from the processor they need only have the processor check the l2 cache in cache hierarchies which do not enforce inclusion the l1 cache must be checked as well as a drawback there is a correlation between the associativities of l1 and l2 caches if the l2 cache does not have at least as many ways as all l1 caches together the effective associativity of the l1 caches is restricted another disadvantage of inclusive cache is that whenever there is an eviction in l2 cache the possibly corresponding lines in l1 also have to get evicted in order to maintain inclusiveness this is quite a bit of work and would result in a higher l1 miss rate 59 another advantage of inclusive caches is that the larger cache can use larger cache lines which reduces the size of the secondary cache tags exclusive caches require both caches to have the same size cache lines so that cache lines can be swapped on a l1 miss l2 hit if the secondary cache is an order of magnitude larger than the primary and the cache data are an order of magnitude larger than the cache tags this tag area saved can be comparable to the incremental area needed to store the l1 cache data in the l2 60 scratchpad memory edit main article scratchpad memory scratchpad memory spm also known as scratchpad scratchpad ram or local store in computer terminology is a high speed internal memory used for temporary storage of calculations data and other work in progress example the k8 edit to illustrate both specialization and multi level caching here is the cache hierarchy of the k8 core in the amd athlon 64 cpu 61 cache hierarchy of the k8 core in the amd athlon 64 cpu the k8 has four specialized caches an instruction cache an instruction tlb a data tlb and a data cache each of these caches is specialized the instruction cache keeps copies of 64 byte lines of memory and fetches 16 bytes each cycle each byte in this cache is stored in ten bits rather than eight with the extra bits marking the boundaries of instructions this is an example of predecoding the cache has only parity protection rather than ecc because parity is smaller and any damaged data can be replaced by fresh data fetched from memory which always has an up to date copy of instructions the instruction tlb keeps copies of page table entries ptes each cycle s instruction fetch has its virtual address translated through this tlb into a physical address each entry is either four or eight bytes in memory because the k8 has a variable page size each of the tlbs is split into two sections one to keep ptes that map 4 kib pages and one to keep ptes that map 4 mib or 2 mib pages the split allows the fully associative match circuitry in each section to be simpler the operating system maps different sections of the virtual address space with different size ptes the data tlb has two copies which keep identical entries the two copies allow two data accesses per cycle to translate virtual addresses to physical addresses like the instruction tlb this tlb is split into two kinds of entries the data cache keeps copies of 64 byte lines of memory it is split into 8 banks each storing 8 kib of data and can fetch two 8 byte data each cycle so long as those data are in different banks there are two copies of the tags because each 64 byte line is spread among all eight banks each tag copy handles one of the two accesses per cycle the k8 also has multiple level caches there are second level instruction and data tlbs which store only ptes mapping 4 kib both instruction and data caches and the various tlbs can fill from the large unified l2 cache this cache is exclusive to both the l1 instruction and data caches which means that any 8 byte line can only be in one of the l1 instruction cache the l1 data cache or the l2 cache it is however possible for a line in the data cache to have a pte which is also in one of the tlbs the operating system is responsible for keeping the tlbs coherent by flushing portions of them when the page tables in memory are updated the k8 also caches information that is never stored in memory prediction information these caches are not shown in the above diagram as is usual for this class of cpu the k8 has fairly complex branch prediction with tables that help predict whether branches are taken and other tables which predict the targets of branches and jumps some of this information is associated with instructions in both the level 1 instruction cache and the unified secondary cache the k8 uses an interesting trick to store prediction information with instructions in the secondary cache lines in the secondary cache are protected from accidental data corruption e g by an alpha particle strike by either ecc or parity depending on whether those lines were evicted from the data or instruction primary caches since the parity code takes fewer bits than the ecc code lines from the instruction cache have a few spare bits these bits are used to cache branch prediction information associated with those instructions the net result is that the branch predictor has a larger effective history table and so has better accuracy more hierarchies edit other processors have other kinds of predictors e g the store to load bypass predictor in the dec alpha 21264 these predictors are caches in that they store information that is costly to compute some of the terminology used when discussing predictors is the same as that for caches one speaks of a hit in a branch predictor but predictors are not generally thought of as part of the cache hierarchy the k8 keeps the instruction and data caches coherent in hardware which means that a store into an instruction closely following the store instruction will change that following instruction other processors like those in the alpha and mips family have relied on software to keep the instruction cache coherent stores are not guaranteed to show up in the instruction stream until a program calls an operating system facility to ensure coherency tag ram edit tag ram on board of a intel pentium iii in computer engineering a tag ram is used to specify which of the possible memory locations is currently stored in a cpu cache 62 63 for a simple direct mapped design fast sram can be used higher associative caches usually employ content addressable memory implementation edit main article cache algorithms cache reads are the most common cpu operation that takes more than a single cycle program execution time tends to be very sensitive to the latency of a level 1 data cache hit a great deal of design effort and often power and silicon area are expended making the caches as fast as possible the simplest cache is a virtually indexed direct mapped cache the virtual address is calculated with an adder the relevant portion of the address extracted and used to index an sram which returns the loaded data the data are byte aligned in a byte shifter and from there are bypassed to the next operation there is no need for any tag checking in the inner loop in fact the tags need not even be read later in the pipeline but before the load instruction is retired the tag for the loaded data must be read and checked against the virtual address to make sure there was a cache hit on a miss the cache is updated with the requested cache line and the pipeline is restarted an associative cache is more complicated because some form of tag must be read to determine which entry of the cache to select an n way set associative level 1 cache usually reads all n possible tags and n data in parallel and then chooses the data associated with the matching tag level 2 caches sometimes save power by reading the tags first so that only one data element is read from the data sram read path for a 2 way associative cache the adjacent diagram is intended to clarify the manner in which the various fields of the address are used address bit 31 is most significant bit 0 is least significant the diagram shows the srams indexing and multiplexing for a 4 kib 2 way set associative virtually indexed and virtually tagged cache with 64 byte b lines a 32 bit read width and 32 bit virtual address because the cache is 4 kib and has 64 b lines there are just 64 lines in the cache and we read two at a time from a tag sram which has 32 rows each with a pair of 21 bit tags although any function of virtual address bits 31 through 6 could be used to index the tag and data srams it is simplest to use the least significant bits similarly because the cache is 4 kib and has a 4 b read path and reads two ways for each access the data sram is 512 rows by 8 bytes wide a more modern cache might be 16 kib 4 way set associative virtually indexed virtually hinted and physically tagged with 32 b lines 32 bit read width and 36 bit physical addresses the read path recurrence for such a cache looks very similar to the path above instead of tags virtual hints are read and matched against a subset of the virtual address later on in the pipeline the virtual address is translated into a physical address by the tlb and the physical tag is read just one as the virtual hint supplies which way of the cache to read finally the physical address is compared to the physical tag to determine if a hit has occurred some sparc designs have improved the speed of their l1 caches by a few gate delays by collapsing the virtual address adder into the sram decoders see sum addressed decoder history edit the early history of cache technology is closely tied to the invention and use of virtual memory citation needed because of scarcity and cost of semi conductor memories early mainframe computers in the 1960s used a complex hierarchy of physical memory mapped onto a flat virtual memory space used by programs the memory technologies would span semi conductor magnetic core drum and disc virtual memory seen and used by programs would be flat and caching would be used to fetch data and instructions into the fastest memory ahead of processor access extensive studies were done to optimize the cache sizes optimal values were found to depend greatly on the programming language used with algol needing the smallest and fortran and cobol needing the largest cache sizes disputed discuss in the early days of microcomputer technology memory access was only slightly slower than register access but since the 1980s 64 the performance gap between processor and memory has been growing microprocessors have advanced much faster than memory especially in terms of their operating frequency so memory became a performance bottleneck while it was technically possible to have all the main memory as fast as the cpu a more economically viable path has been taken use plenty of low speed memory but also introduce a small high speed cache memory to alleviate the performance gap this provided an order of magnitude more capacity for the same price with only a slightly reduced combined performance first tlb implementations edit the first documented uses of a tlb were on the ge 645 65 and the ibm 360 67 66 both of which used an associative memory as a tlb first instruction cache edit the first documented use of an instruction cache was on the cdc 6600 67 first data cache edit the first documented use of a data cache was on the ibm system 360 model 85 68 in 68k...
|