Meta tags:
Headings (most frequently used words):
cache, in, policies, associative, multi, first, microprocessors, cpu, history, operation, two, way, example, and, caches, write, instruction, versus, contents, overview, associativity, entry, structure, miss, address, translation, hierarchy, modern, processor, implementation, see, also, notes, references, external, links, entries, performance, direct, mapped, set, speculative, execution, skewed, pseudo, multicolumn, flag, bits, homonym, synonym, problems, virtual, tags, hints, page, coloring, specialized, level, scratchpad, memory, the, k8, more, hierarchies, tag, ram, ported, replacement, stalls, victim, trace, coalescing, wcc, micro, μop, or, uop, branch, target, smart, core, chips, separate, unified, exclusive, inclusive, tlb, implementations, data, 68k, x86, arm, current, research,
Text of the page (most frequently used words):
the (933), cache (604), and (288), #memory (176), for (145), data (125), caches (110), with (107), that (103), from (100), address (90), virtual (89), are (84), which (79), #instruction (77), this (70), have (67), can (66), cpu (62), associative (62), processor (60), main (60), physical (59), way (58), edit (58), one (58), used (57), level (55), has (54), set (53), not (49), intel (48), tag (45), core (44), performance (43), two (42), each (41), bits (39), bit (38), instructions (37), access (36), retrieved (35), tlb (35), also (35), read (35), system (34), index (33), pdf (32), kib (32), may (31), was (31), chip (31), write (31), there (31), some (31), page (30), entry (30), but (30), into (28), time (28), location (28), more (28), than (28), use (27), computer (27), size (27), processors (27), only (27), tags (27), other (26), when (26), ibm (25), mapped (25), first (25), miss (25), line (25), all (24), different (24), unit (24), multi (24), larger (24), same (23), pages (23), block (23), high (22), direct (22), power (21), architecture (21), been (21), original (21), speed (21), addresses (21), mib (21), both (21), blocks (21), hierarchy (20), archived (20), number (20), per (19), between (19), latency (19), because (19), however (19), hardware (18), its (18), associativity (18), micro (18), most (18), any (18), lines (18), hit (18), cores (18), cpus (17), branch (17), translation (17), shared (17), they (17), multiple (17), design (16), register (16), store (16), doi (16), operation (16), amd (16), were (16), much (16), virtually (16), then (16), locations (16), program (15), systems (15), article (15), sram (15), does (15), fetch (15), indexed (15), entries (15), misses (15), policy (14), processing (14), four (14), possible (14), space (14), physically (14), such (14), usually (14), generally (14), called (14), replacement (13), policies (13), single (13), execution (13), machine (13), another (13), part (13), 2013 (13), operating (13), had (13), levels (13), early (13), these (13), faster (13), byte (13), since (13), tagged (13), must (13), will (13), mapping (13), buffer (12), pipeline (12), stored (12), based (12), trace (12), pentium (12), large (12), example (12), least (12), before (12), those (12), table (11), value (11), rate (11), 128 (11), microarchitecture (11), caching (11), very (11), inclusive (11), well (11), fast (11), see (11), where (11), area (11), microprocessors (11), die (11), motherboard (11), split (11), need (11), separate (11), cached (11), μop (11), available (10), microprocessor (10), management (10), mmu (10), low (10), order (10), 360 (10), arm (10), conflict (10), 2012 (10), 2014 (10), small (10), new (10), major (10), would (10), ram (10), through (10), sizes (10), victim (10), after (10), eight (10), modern (10), having (10), thus (10), implemented (10), decoded (10), using (9), non (9), model (9), history (9), 256 (9), several (9), times (9), international (9), 1990 (9), exclusive (9), information (9), case (9), designs (9), advantage (9), slower (9), either (9), above (9), typically (9), ways (9), keep (9), copies (9), copy (9), smaller (9), require (9), many (9), toggle (8), computing (8), control (8), target (8), bus (8), scratchpad (8), cycle (8), secondary (8), 2009 (8), symposium (8), software (8), work (8), cost (8), skewed (8), stores (8), fetched (8), edram (8), common (8), their (8), hints (8), hint (8), although (8), simple (8), like (8), unified (8), accesses (8), rather (8), specialized (8), include (8), coloring (8), problem (8), multicolumn (8), contents (7), wikipedia (7), cs1 (7), 2010 (7), 2008 (7), deprecated (7), service (7), general (7), dynamic (7), application (7), package (7), types (7), second (7), x86 (7), external (7), 1109 (7), proceedings (7), conference (7), 2020 (7), 1145 (7), problems (7), isbn (7), 2015 (7), cite (7), anandtech (7), 2001 (7), length (7), september (7), uses (7), given (7), allows (7), important (7), just (7), dram (7), three (7), bytes (7), while (7), path (7), few (7), function (7), evicted (7), fewer (7), known (7), become (7), color (7), colors (7), tagging (7), requires (7), offset (7), valid (7), flag (7), code (6), last (6), december (6), 2023 (6), org (6), module (6), clock (6), integrated (6), decoder (6), registers (6), predictor (6), load (6), coherence (6), graphics (6), operations (6), cycles (6), prediction (6), random (6), written (6), november (6), ieee (6), s2cid (6), smart (6), haswell (6), fully (6), gap (6), tlbs (6), help (6), total (6), placement (6), effectively (6), recently (6), still (6), significant (6), form (6), extra (6), taken (6), programs (6), needed (6), later (6), determine (6), requested (6), often (6), until (6), whether (6), among (6), map (6), simpler (6), structure (6), run (6), longer (6), hold (6), makes (6), delay (6), particular (6), selected (6), lower (6), below (6), back (6), hash (6), dirty (6), log (6), due (6), subsection (6), search (5), links (5), articles (5), march (5), maint (5), archival (5), central (5), quantum (5), file (5), logic (5), lookaside (5), variable (5), word (5), thread (5), wide (5), issue (5), out (5), risc (5), alpha (5), mips (5), sets (5), turing (5), what (5), 2003 (5), family (5), solutions (5), parallel (5), com (5), introduction (5), link (5), pro (5), 978 (5), bridge (5), atlas (5), over (5), technology (5), reduced (5), llc (5), fetching (5), instead (5), added (5), iii (5), board (5), next (5), implementation (5), being (5), rates (5), could (5), improve (5), right (5), released (5), following (5), provided (5), done (5), various (5), associated (5), checking (5), checked (5), make (5), takes (5), means (5), result (5), better (5), ptes (5), storing (5), allow (5), match (5), get (5), patterns (5), full (5), needs (5), reduce (5), aliases (5), existing (5), vipt (5), contains (5), pseudo (5), evict (5), about (4), additional (4), terms (4), august (4), statements (4), short (4), security (4), electronics (4), purpose (4), signal (4), functional (4), magnitude (4), pipelined (4), speculative (4), process (4), scalar (4), isa (4), itanium (4), sparc (4), series (4), automaton (4), technologies (4), flushing (4), 2004 (4), reference (4), acm (4), xeon (4), 1968 (4), october (4), cortex (4), 2007 (4), hierarchies (4), advanced (4), techniques (4), web (4), iris (4), 2011 (4), 2016 (4), frontend (4), david (4), decode (4), uop (4), optimization (4), zhang (4), slice (4), scheme (4), predicting (4), fact (4), approach (4), works (4), tables (4), every (4), ported (4), accessing (4), benefit (4), tools (4), energy (4), recent (4), current (4), depending (4), type (4), unusually (4), introduced (4), dynamically (4), became (4), included (4), directly (4), chips (4), loop (4), special (4), implementations (4), complex (4), fastest (4), might (4), similar (4), against (4), compared (4), reads (4), diagram (4), indexing (4), reading (4), higher (4), content (4), keeps (4), ensure (4), coherency (4), predictors (4), kinds (4), parity (4), effective (4), maps (4), athlon (4), comparable (4), hits (4), versus (4), highest (4), increasing (4), accessed (4), reducing (4), complexity (4), protocols (4), amount (4), off (4), examples (4), actual (4), without (4), avoiding (4), wcc (4), writes (4), already (4), guarantee (4), alternatively (4), changes (4), mappings (4), difficult (4), within (4), differences (4), best (4), aliasing (4), less (4), hide (4), move (4), sidebar (4), mobile (3), view (3), organization (3), long (3), unsourced (3), array (3), digital (3), related (3), frequency (3), switch (3), sum (3), addressed (3), adder (3), counter (3), sequential (3), controller (3), point (3), generation (3), others (3), 512 (3), network (3), gpu (3), multiprocessor (3), embedded (3), vector (3), updates (3), simultaneous (3), algorithm (3), false (3), sharing (3), stall (3), dec (3), comparison (3), edge (3), specific (3), storage (3), cellular (3), machines (3), queue (3), review (3), updated (3), simulation (3), smith (3), introduces (3), capacity (3), annual (3), product (3), world (3), james (3), overview (3), 486 (3), corporation (3), microcomputer (3), 2nd (3), characteristics (3), joint (3), third (3), february (3), 645 (3), manual (3), technical (3), analysis (3), usa (3), jouppi (3), norman (3), sigarch (3), news (3), tested (3), 2006 (3), april (3), 2018 (3), how (3), real (3), bulldozer (3), programmers (3), compiler (3), buffers (3), 2024 (3), lake (3), multicore (3), computers (3), john (3), selection (3), shown (3), advantages (3), conventional (3), consumption (3), atom (3), 1976 (3), memories (3), intended (3), idea (3), placed (3), loops (3), 2021 (3), simulator (3), algorithms (3), request (3), efficiency (3), apple (3), feature (3), uncommon (3), find (3), took (3), slightly (3), mhz (3), located (3), devices (3), 68k (3), documented (3), especially (3), found (3), smallest (3), largest (3), subset (3), occurred (3), rows (3), pair (3), simplest (3), select (3), sometimes (3), loaded (3), even (3), making (3), addressable (3), coherent (3), ecc (3), predict (3), branches (3), them (3), banks (3), sections (3), ten (3), always (3), local (3), instance (3), intermediate (3), present (3), certain (3), should (3), cray (3), itself (3), down (3), l1i (3), holds (3), providing (3), required (3), traces (3), fixed (3), decreasing (3), coalescing (3), points (3), technique (3), during (3), likely (3), pattern (3), lead (3), necessary (3), described (3), flushed (3), changed (3), becomes (3), solved (3), writing (3), synonym (3), homonym (3), historically (3), lookup (3), choice (3), r6000 (3), looked (3), vivt (3), refer (3), context (3), implement (3), comparing (3), cause (3), continue (3), invalid (3), stale (3), row (3), closer (3), lru (3), transistors (3), places (3), doubling (3), choose (3), waiting (3), stalls (3), checks (3), exception (3), settle (3), topic (2), languages (2), contact (2), privacy (2), wikimedia (2), commons (2), categories (2), wayback (2), disputed (2), accuracy (2), description (2), wikidata (2), pin (2), semiconductor (2), device (2), ppw (2), watt (2), voltage (2), scaling (2), circuit (2), circuitry (2), barrel (2), shifter (2), datapath (2), stack (2), gate (2), floating (2), fifo (2), heterogeneous (2), count (2), slicing (2), image (2), coprocessor (2), stream (2), transistor (2), temporal (2), multithreading (2), task (2), pipelining (2), parallelism (2), renaming (2), structural (2), hazards (2), classic (2), visc (2), motorola (2), addressing (2), architectures (2), finite (2), state (2), models (2), lecture (2), power4 (2), understanding (2), introductory (2), 2002 (2), hill (2), paper (2), results (2), benchmarks (2), net (2), describing (2), cacti (2), foster (2), chen (2), june (2), a22 (2), 6916 (2), 1964 (2), 6600 (2), 1999 (2), bottleneck (2), 123 (2), hotchips (2), cluster (2), tian (2), 5200 (2), 4950hq (2), difference (2), leon (2), cutress (2), ian (2), zen (2), shimpi (2), anand (2), lal (2), islped (2), aware (2), kanter (2), sandy (2), continued (2), agner (2), guide (2), ghz (2), skylake (2), desktop (2), products (2), formerly (2), improving (2), 17th (2), isca (2), jiang (2), lin (2), xiaodong (2), george (2), mechanism (2), journal (2), citeseerx (2), vol (2), 1962 (2), ifip (2), congress (2), patterson (2), hennessy (2), interface (2), choosing (2), error (2), protection (2), patents (2), justia (2), patent (2), issued (2), 1994 (2), reconfigurable (2), 1997 (2), schemes (2), ryan (2), codenamed (2), ivy (2), networks (2), z13 (2), throughput (2), 1982 (2), developed (2), cambridge (2), sequence (2), slave (2), give (2), effect (2), communication (2), references (2), notes (2), allocation (2), locality (2), serve (2), traditional (2), normally (2), whereas (2), connected (2), entirely (2), average (2), consider (2), research (2), 192 (2), laptop (2), variant (2), serves (2), crystalwell (2), again (2), tens (2), popularity (2), socket (2), obsolete (2), began (2), onto (2), sockets (2), made (2), growing (2), versions (2), 386 (2), amounts (2), expensive (2), around (2), left (2), 68040 (2), 68020 (2), mode (2), considered (2), tiny (2), replaced (2), true (2), 1980s (2), introduce (2), price (2), closely (2), semi (2), conductor (2), mainframe (2), 1960s (2), flat (2), extensive (2), values (2), depend (2), greatly (2), needing (2), citation (2), decoders (2), hinted (2), width (2), recurrence (2), translated (2), finally (2), srams (2), shows (2), matching (2), calculated (2), portion (2), retired (2), tends (2), sensitive (2), great (2), silicon (2), currently (2), change (2), guaranteed (2), terminology (2), primary (2), never (2), fill (2), keeping (2), translate (2), marking (2), date (2), specialization (2), here (2), reduces (2), strictly (2), check (2), together (2), disadvantage (2), whenever (2), corresponding (2), quite (2), name (2), separately (2), meaning (2), benefits (2), including (2), ranges (2), independently (2), resulting (2), reasons (2), allowing (2), share (2), question (2), implementing (2), files (2), did (2), end (2), scheduled (2), allocates (2), nest (2), phenom (2), features (2), along (2), returned (2), cases (2), replacing (2), fundamental (2), tradeoff (2), operate (2), dedicated (2), equal (2), method (2), powered (2), delivering (2), enough (2), causing (2), created (2), heuristic (2), l1d (2), basic (2), extreme (2), commonly (2), property (2), role (2), server (2), probes (2), mentioned (2), 6000 (2), attempt (2), assign (2), simply (2), extremely (2), collide (2), arrange (2), literature (2), community (2), cannot (2), consistent (2), timing (2), perhaps (2), suffers (2), determining (2), alternate (2), matches (2), distinguish (2), simultaneously (2), perfect (2), proceed (2), allowed (2), inconsistent (2), overlapping (2), spaces (2), now (2), tend (2), pipt (2), homonyms (2), designed (2), exist (2), potential (2), additionally (2), avoids (2), slow (2), divided (2), held (2), running (2), cut (2), dependent (2), measurement (2), indicates (2), marked (2), implies (2), field (2), hence (2), displaystyle (2), lceil (2), rceil (2), ratio (2), contributes (2), conflicts (2), tests (2), something (2), good (2), suffer (2), hand (2), unpredictable (2), acts (2), free (2), place (2), execute (2), independent (2), waits (2), hundreds (2), able (2), steps (2), immediately (2), rarely (2), future (2), contain (2), copied (2), etc (2), static (2), executable (2), step (2), appearance (2), upload (2), create (2), account (2), donate (2), menu (2), add, cookie, statement, statistics, developers, conduct, legal, safety, contacts, disclaimers, text, under, apply, site, you, agree, registered, trademark, profit, foundation, inc, creative, attribution, sharealike, license, rendered, parsoid, edited, 2026, utc, hidden, webarchive, template, volume, disputes, errors, parameters, https, php, title, cpu_cache, oldid, 1372368516, carrier, grid, tick, tock, fabrication, chronology, gating, acpi, apm, pmu, analog, boolean, mixed, binary, multiplier, demultiplexer, multiplexer, rom, microcode, hardwired, status, glue, combinational, imc, fpu, agu, alu, arithmetic, units, components, manycore, baseband, secure, cryptoprocessor, tpu, tensor, dsp, ppu, physics, vpu, vision, accelerator, accelerators, noc, cypress, psoc, mpsoc, soc, soft, asip, ultra, microcontroller, pop, sip, mcm, cpld, fpoa, fpga, asic, pal, tile, gpgpu, orders, metrics, sups, synaptic, tps, transactions, flops, ips, ipc, cpi, spmd, mimd, misd, swar, simt, simd, sisd, flynn, taxonomy, cooperative, preemptive, heterogenous, hyperthreading, distributed, superscalar, serial, dependence, reservation, station, tomasulo, scoreboarding, dependency, operand, forwarding, epiphany, tilera, 3x0, 390, 370, lmc, microblaze, openrisc, unicore, m32r, etrax, cris, superh, clipper, powerpc, stanford, pdp, vax, 68000, modes, zisc, nisc, oisc, misc, epic, vliw, trips, cisc, orthogonal, neuromorphic, cognitive, multiprocessing, fabric, huma, numa, endianness, transport, triggered, dataflow, modified, harvard, von, neumann, pointer, belt, zeno, hypercomputation, probabilistic, nondeterministic, post, universal, alternating, deterministic, hierarchical, abstract, fallacy, princeton, university, ixbtlabs, pavel, danilov, ars, technica, jon, stokes, vhdl, paul, genua, freescale, primer, ruud, van, der, pas, sun, microsystems, cantin, thorough, lucidly, presented, reasonably, organizations, spec, cpu2000, 1989, compulsory, classification, evaluating, lwn, ulrich, drepper, detail, wikibook, labs, wang, zhenghong, lee, ruby, 41st, novel, enhanced, sally, adee, 43892134, mspec, 5292036, spectrum, thwarts, sneak, attack, ark, specs, july, reilly, kheradpir, shervin, allan, flight, thornton, fall, edition, 1972, ga27, 2719, january, electric, mahapatra, nihar, venkatrao, balakrishna, 3es, 11557476, 357783, 331677, crossroads, l210, r4f, sandpile, jaleel, aamer, eric, borch, bhandaru, malini, steely, simon, emer, joel, jaleels, achieving, zheng, ying, davis, brian, jordan, matthew, austin, texas, 7803, 8385, ispass, 1291359, evaluation, amecomputers, explanation, bradley, borg, anita, 1992, 114, 146628, 139708, study, lempel, oded, cornell, workshop, interconnects, shih, chiu, edaboard, betw, inside, demo, niu, kun, btic, motiani, dipti, 2017, dual, schedulers, revealed, analyzed, solomon, baruch, mendelson, avi, orenstein, doron, almog, yoav, ronen, ronny, huntington, beach, 195859085, 58113, 371, lpe, 945363, association, machinery, subsystem, via, assembly, makers, fog, 2000, launch, crystal, addition, prefetch, seattle, 373, 134547, 364, letter, qingda, ding, xiaoning, zhao, sadayappan, 14th, salt, city, utah, 378, hpca, 4658653, 367, gaining, insights, partitioning, bridging, roscoe, timothy, baumann, andrew, ethz, 263, 3800, 00l, taylor, davies, peter, farmwald, michael, 2si, 363, 325096, 325161, 355, bottomley, linux, parameter, kaxiras, stefanos, ros, alberto, 40th, 547, 15434231, 9781450320795, 2485922, 2485968, 307, 9125, 535, perspective, efficient, kilburn, payne, howarth, 1961, conferences, eastern, washington, macmillan, 294, 279, key, supervisor, sumner, haley, chenh, spartan, neill, proc, afips, spring, 1967, 621, 1465482, 1465581, 611, experience, multiprogramming, relocation, cragon, harvey, 1996, jones, bartlett, learning, 209, 86720, 474, dugan, ben, concerning, cooperman, gene, basics, morgan, kaufmann, 484, 374493, sadler, nathan, sorin, daniel, 20160350229, developer, programmer, chenxi, yan, yong, 621212, 1997imicr, 17e, 40c, bibcode, kozyrakis, seznec, andré, 1993, 178, 173682, 165152, 169, jahagirdar, sanjeev, varghese, sodhi, inder, wells, megalingam, rajesh, kannan, deepu, joseph, iype, vikram, vandana, science, 556, 18236635, 4244, 4519, iccsit, 5234663, 551, phased, ucsd, edu, launches, p5900, 10nm, radio, press, release, 32kb, 5mb, 15mb, newsroom, sheet, accelerating, infrastructure, white, bill, cecilia, z13s, elsevier, 383872, quantitative, altering, raise, suggest, researchers, alan, jay, 530, 6023466, 356887, 356892, 473, surveys, liptay, 1147, 0015, aspects, landy, barry, tunnel, diode, worked, speeded, operands, obeyed, modulo, remaining, wanted, speedup, words, mathematical, laboratory, aldermaston, cad, centre, chao, zeng, qingkai, nicopolitidis, petros, 1939, 0122, issn, 1155, 5559552, survey, side, channel, attacks, systematic, countermeasures, torres, gabriel, paging, frame, ferranti, memoization, dinero, prefetching, ports, phases, concept, super, architects, explore, tradeoffs, simplescalar, open, source, options, focused, fault, tolerance, goals, mainframes, equipped, gt3e, increasingly, newer, generations, megabytes, enjoyed, prolonged, thanks, previously, named, motherboards, produced, disappeared, development, brought, clocked, termed, differentiate, boards, contained, 485turbocache, kbyte, era, disparity, caused, sdram, mmx, daughtercard, support, reached, featured, 120, refresh, constructed, significantly, latencies, nine, enable, optional, upgrade, dip, cells, bottom, corner, a38202, austek, simms, i386, 68060, kilobytes, 1987, basically, shrink, burst, 68030, accelerates, consist, 1984, typical, 68010, cdc, days, technically, economically, viable, plenty, alleviate, combined, tied, invention, scarcity, span, magnetic, drum, disc, seen, ahead, studies, optimize, optimal, programming, language, algol, fortran, cobol, discuss, improved, delays, collapsing, looks, matched, supplies, similarly, adjacent, clarify, manner, fields, multiplexing, complicated, chooses, save, element, relevant, extracted, returns, aligned, bypassed, inner, sure, restarted, deal, effort, expended, engineering, specify, employ, relied, show, calls, facility, costly, compute, discussing, speaks, thought, bypass, 21264, interesting, trick, protected, accidental, corruption, strike, spare, particle, usual, class, fairly, targets, jumps, pte, responsible, portions, spread, handles, identical, section, fetches, boundaries, predecoding, damaged, fresh, illustrate, spm, internal, temporary, calculations, progress, swapped, saved, incremental, wish, remove, enforce, inclusion, drawback, correlation, associativities, restricted, eviction, possibly, maintain, inclusiveness, diminishes, hitting, exchanged, exchange, copying, decisions, somewhere, reside, universally, accepted, names, partially, demonstrated, constraint, referred, pieces, undesirable, increase, considerably, global, desirable, whole, redundancy, processes, threads, utilized, considering, inevitably, wiring, circa, usable, characteristic, assignments, reallocated, runtime, bank, break, dependencies, easing, acting, essentially, tulsa, incorporated, compatible, madison, 1995, 21164, incorporating, begun, utilize, pull, entire, 2010s, mounted, fourth, rare, 2019, flip, flop, starting, oryon, qualcomm, lunar, a14, z15
Text of the page (random words):
ore power and chip area and potentially more time on the other hand caches with more associativity suffer fewer misses see conflict misses so that the cpu wastes less time reading from the slow main memory the general guideline is that doubling the associativity from direct mapped to two way or from two way to four way has about the same effect on raising the hit rate as doubling the cache size however increasing associativity more than four does not improve hit rate as much 13 and are generally done for other reasons see virtual aliasing some cpus can dynamically reduce the associativity of their caches in low power states which acts as a power saving measure 14 in order of worse but simple to better but complex direct mapped cache good best case time but unpredictable in the worst case two way set associative cache two way skewed associative cache 15 four way set associative cache eight way set associative cache a common choice for later implementations 12 way set associative cache similar to eight way fully associative cache the best miss rates but practical only for a small number of entries direct mapped cache edit in this cache organization each location in the main memory can go in only one entry in the cache therefore a direct mapped cache can also be called a one way set associative cache it does not have a placement policy as such since there is no choice of which cache entry s contents to evict this means that if two locations map to the same entry they may continually knock each other out although simpler a direct mapped cache needs to be much larger than an associative one to give comparable performance and it is more unpredictable let x be block number in cache y be block number of memory and n be number of blocks in cache then mapping is done with the help of the equation x y mod n two way set associative cache edit if each location in the main memory can be cached in either of two locations in the cache one logical question is which one of the two the simplest and most commonly used scheme shown in the right hand diagram above is to use the least significant bits of the memory location s index as the index for the cache memory and to have two entries for each index one benefit of this scheme is that the tags stored in the cache do not have to include that part of the main memory address which is implied by the cache memory s index since the cache tags have fewer bits they require fewer transistors take less space on the processor circuit board or on the microprocessor chip and can be read and compared faster also lru algorithm is especially simple since only one bit needs to be stored for each pair speculative execution edit one of the advantages of a direct mapped cache is that it allows simple and fast speculation once the address has been computed the one cache index which might have a copy of that location in memory is known that cache entry can be read and the processor can continue to work with that data before it finishes checking that the tag actually matches the requested address the idea of having the processor use the cached data before the tag match completes can be applied to associative caches as well a subset of the tag called a hint can be used to pick just one of the possible cache entries mapping to the requested address the entry selected by the hint can then be used in parallel with checking the full tag the hint technique works best when used in the context of address translation as explained below two way skewed associative cache edit other schemes have been suggested such as the skewed cache 15 where the index for way 0 is direct as above but the index for way 1 is formed with a hash function a good hash function has the property that addresses which conflict with the direct mapping tend not to conflict when mapped with the hash function and so it is less likely that a program will suffer from an unexpectedly large number of conflict misses due to a pathological access pattern the downside is extra latency from computing the hash function 16 additionally when it comes time to load a new line and evict an old line it may be difficult to determine which existing line was least recently used because the new line conflicts with data at different indexes in each way lru tracking for non skewed caches is usually done on a per set basis nevertheless skewed associative caches have major advantages over conventional set associative ones 17 pseudo associative cache edit a true set associative cache tests all the possible ways simultaneously using something like a content addressable memory a pseudo associative cache tests each possible way one at a time a hash rehash cache and a column associative cache are examples of a pseudo associative cache in the common case of finding a hit in the first way tested a pseudo associative cache is as fast as a direct mapped cache but it has a much lower conflict miss rate than a direct mapped cache closer to the miss rate of a fully associative cache 16 multicolumn cache edit comparing with a direct mapped cache a set associative cache has a reduced number of bits for its cache set index that maps to a cache set where multiple ways or blocks stays such as 2 blocks for a 2 way set associative cache and 4 blocks for a 4 way set associative cache comparing with a direct mapped cache the unused cache index bits become a part of the tag bits for example a 2 way set associative cache contributes 1 bit to the tag and a 4 way set associative cache contributes 2 bits to the tag the basic idea of the multicolumn cache 18 is to use the set index to map to a cache set as a conventional set associative cache does and to use the added tag bits to index a way in the set for example in a 4 way set associative cache the two bits are used to index way 00 way 01 way 10 and way 11 respectively this double cache indexing is called a major location mapping and its latency is equivalent to a direct mapped access extensive experiments in multicolumn cache design 18 shows that the hit ratio to major locations is as high as 90 if cache mapping conflicts with a cache block in the major location the existing cache block will be moved to another cache way in the same set which is called selected location because the newly indexed cache block is a most recently used mru block it is placed in the major location in multicolumn cache with a consideration of temporal locality since multicolumn cache is designed for a cache with a high associativity the number of ways in each set is high thus it is easy find a selected location in the set a selected location index by an additional hardware is maintained for the major location in a cache block citation needed multicolumn cache remains a high hit ratio due to its high associativity and has a comparable low latency to a direct mapped cache due to its high percentage of hits in major locations the concepts of major locations and selected locations in multicolumn cache have been used in several cache designs in arm cortex r chip 19 intel s way predicting cache memory 20 ibm s reconfigurable multi way associative cache memory 21 and oracle s dynamic cache replacement way selection based on address tab bits 22 cache entry structure edit cache row entries usually have the following structure tag data block flag bits the data block cache line contains the actual data fetched from the main memory the tag contains part of the address of the actual data fetched from the main memory the flag bits are discussed below the size of the cache is the amount of main memory data it can hold this size can be calculated as the number of bytes stored in each data block times the number of blocks stored in the cache the tag flag and error correction code bits are not included in the size 23 although they do affect the physical area of a cache an effective memory address which goes along with the cache line memory block is split msb to lsb into the tag the index and the block offset 8 24 tag index block offset the index describes which cache set that the data has been put in the index length is log 2 s displaystyle lceil log _ 2 s rceil bits for s cache sets the block offset specifies the desired data within the stored data block within the cache row typically the effective address is in bytes so the block offset length is log 2 b displaystyle lceil log _ 2 b rceil bits where b is the number of bytes per data block the tag contains the most significant bits of the address which are checked against all rows in the current set the set has been retrieved by index to see if this set contains the requested address if it does a cache hit occurs the tag length in bits is as follows tag_length address_length index_length block_offset_length some authors refer to the block offset as simply the offset 25 or the displacement 26 27 example edit the original pentium 4 processor had a four way set associative l1 data cache of 8 kib in size with 64 byte cache blocks hence there are 8 kib 64 128 cache blocks the number of sets is equal to the number of cache blocks divided by the number of ways of associativity what leads to 128 4 32 sets and hence 2 5 32 different indices there are 2 6 64 possible offsets since the cpu address is 32 bits wide this implies 32 5 6 21 bits for the tag field the original pentium 4 processor also had an eight way set associative l2 integrated cache 256 kib in size with 128 byte cache blocks this implies 32 8 7 17 bits for the tag field 25 flag bits edit an instruction cache requires only one flag bit per cache row entry a valid bit the valid bit indicates whether or not a cache block has been loaded with valid data on power up the hardware sets all the valid bits in all the caches to invalid some systems also set a valid bit to invalid at other times such as when multi master bus snooping hardware in the cache of one processor hears an address broadcast from some other processor and realizes that certain data blocks in the local cache are now stale and should be marked invalid a data cache typically requires two flag bits per cache line a valid bit and a dirty bit having a dirty bit set indicates that the associated cache line has been changed since it was read from main memory dirty meaning that the processor has written data to that line and the new value has not propagated all the way to main memory cache miss edit a cache miss is a failed attempt to read or write a piece of data in the cache which results in a main memory access with much longer latency there are three kinds of cache misses instruction read miss data read miss and data write miss cache read misses from an instruction cache generally cause the largest delay because the processor or at least the thread of execution has to wait stall until the instruction is fetched from main memory cache read misses from a data cache usually cause a smaller delay because instructions not dependent on the cache read can be issued and continue execution until the data are returned from main memory and the dependent instructions can resume execution cache write misses to a data cache generally cause the shortest delay because the write can be queued and there are few limitations on the execution of subsequent instructions the processor can continue until the queue is full for a detailed introduction to the types of misses see cache performance measurement and metric address translation edit most general purpose cpus implement some form of virtual memory to summarize either each program running on the machine sees its own simplified address space which contains code and data for that program only or all programs run in a common virtual address space a program executes by calculating comparing reading and writing to addresses of its virtual address space rather than addresses of physical address space making programs simpler and thus easier to write virtual memory requires the processor to translate virtual addresses generated by the program into physical addresses in main memory the portion of the processor that does this translation is known as the memory management unit mmu the fast path through the mmu can perform those translations stored in the translation lookaside buffer tlb which is a cache of mappings from the operating system s page table segment table or both for the purposes of the present discussion there are three important features of address translation latency the physical address is available from the mmu some time perhaps a few cycles after the virtual address is available from the address generator aliasing multiple virtual addresses can map to a single physical address most processors guarantee that all updates to that single physical address will happen in program order to deliver on that guarantee the processor must ensure that only one copy of a physical address resides in the cache at any given time granularity the virtual address space is broken up into pages for instance a 4 gib virtual address space might be cut up into 1 048 576 pages of 4 kib size each of which can be independently mapped there may be multiple page sizes supported see virtual memory for elaboration one early virtual memory system the ibm m44 44x required an access to a mapping table held in core memory before every programmed access to main memory 28 nb 1 with no caches and with the mapping table memory running at the same speed as main memory this effectively cut the speed of memory access in half two early machines that used a page table in main memory for mapping the ibm system 360 model 67 and the ge 645 both had a small associative memory as a cache for accesses to the in memory page table both machines predated the first machine with a cache for main memory the ibm system 360 model 85 so the first hardware cache used in a computer system was not a data or instruction cache but rather a tlb caches can be divided into four types based on whether the index or tag correspond to physical or virtual addresses physically indexed physically tagged pipt caches use the physical address for both the index and the tag while this is simple and avoids problems with aliasing it is also slow as the physical address must be looked up which could involve a tlb miss and access to main memory before that address can be looked up in the cache virtually indexed virtually tagged vivt caches use the virtual address for both the index and the tag this caching scheme can result in much faster lookups since the mmu does not need to be consulted first to determine the physical address for a given virtual address however vivt suffers from aliasing problems where several different virtual addresses may refer to the same physical address the result is that such addresses would be cached separately despite referring to the same memory causing coherency problems although solutions to this problem exist 31 they do not work for standard coherence protocols another problem is homonyms where the same virtual address maps to seve...
|