Meta tags:
description= Isn t the world 2D?;
Headings (most frequently used words):
the, in, of, cairo, ickle, and, over, red, corner, glamorous, refresh, you, want, threads, why, not, zoidberg, any, old, ion, tale, processors, sandybridge, acceleration, 845g, rises, like, phoenix, from, ashes, xaa, glamor, beginning, going, retro, pages, categories, archives, search, blogroll, rss, feeds, meta,
Text of the page (most frequently used words):
the (280), and (113), for (45), that (44), gpu (36), cairo (31), with (28), performance (27), was (26), from (26), memory (25), using (24), have (23), this (21), than (21), all (20), sna (20), not (20), driver (20), cpu (20), acceleration (19), #glamor (19), uxa (18), slower (18), faster (17), can (17), are (17), more (17), but (16), exa (15), render (14), comments (12), 2012 (12), sandybridge (12), about (11), xaa (11), its (11), software (11), image (11), intel (10), then (10), posted (10), pixmap (10), much (10), copy (10), processor (10), design (9), which (9), you (9), new (9), only (9), also (9), rendering (9), look (9), improvement (9), window (9), very (8), graphics (8), single (8), into (8), cases (8), able (8), threads (8), 2013 (7), few (7), even (7), has (7), those (7), one (7), however (7), should (7), better (7), results (7), being (7), 2500 (7), backend (7), desktop (7), generation (7), get (6), now (6), kernel (6), see (6), years (6), were (6), fast (6), between (6), both (6), any (6), core (6), architecture (6), operations (6), protocol (6), some (6), improved (6), over (6), square (6), rasteriser (6), current (6), wordpress (5), ickle (5), benchmarks (5), take (5), how (5), such (5), could (5), quite (5), without (5), through (5), there (5), what (5), improvements (5), still (5), keep (5), use (5), power (5), test (5), problem (5), increase (5), 2520m (5), here (5), pixman (5), comparing (5), average (5), does (5), ddx (5), like (4), com (4), august (4), december (4), january (4), video (4), been (4), question (4), management (4), where (4), given (4), slow (4), simply (4), cache (4), system (4), operation (4), back (4), having (4), would (4), these (4), did (4), really (4), compare (4), opengl (4), hardware (4), indeed (4), char (4), 10x10 (4), 100x100 (4), 500x500 (4), baseline (4), well (4), just (4), find (4), old (4), run (4), fglrx (4), r600g (4), snb (4), traces (4), compared (4), radeon (4), core2 (4), processors (4), sse2 (4), haswell (4), due (4), nvidia (4), equal (4), existing (4), log (3), 2009 (3), july (3), 2010 (3), october (3), overall (3), longer (3), under (3), ancient (3), speed (3), long (3), surfaces (3), used (3), simple (3), known (3), always (3), running (3), lots (3), execution (3), command (3), buffers (3), gem (3), complete (3), turns (3), out (3), within (3), cost (3), start (3), last (3), level (3), perform (3), migration (3), accelerate (3), efficient (3), need (3), deliver (3), many (3), actually (3), enabling (3), finally (3), least (3), good (3), again (3), upon (3), accelerating (3), time (3), couple (3), ivybridge (3), most (3), putimage (3), shmputimage (3), gpus (3), chipsets (3), never (3), batch (3), small (3), introduction (3), directly (3), xlib (3), despite (3), almost (3), q9550 (3), 3720qm (3), bit (3), same (3), set (3), main (3), mobile (3), 95w (3), they (3), want (3), two (3), chips (3), completely (3), integrated (3), advantage (3), outperform (3), threaded (3), make (3), cores (3), kernels (3), saturate (3), bus (3), isn (3), world (3), site (2), website (2), required (2), write (2), bar (2), view (2), sign (2), subscribed (2), subscribe (2), account (2), create (2), posts (2), november (2), june (2), may (2), categories (2), bugs (2), older (2), large (2), pinch (2), salt (2), geometric (2), mean (2), x11perf (2), change (2), against (2), compiled (2), xf86 (2), things (2), tuning (2), managed (2), ago (2), card (2), allocation (2), server (2), pixmaps (2), downside (2), amount (2), added (2), textures (2), idea (2), needs (2), ask (2), people (2), going (2), real (2), difference (2), since (2), domains (2), igp (2), snoopable (2), domain (2), shared (2), try (2), gtt (2), found (2), various (2), instead (2), made (2), read (2), effectively (2), switch (2), beginning (2), claim (2), improving (2), reach (2), benefit (2), impact (2), move (2), kms (2), drawing (2), reality (2), plenty (2), room (2), perspective (2), lets (2), each (2), seen (2), release (2), library (2), drivers (2), commands (2), ability (2), larger (2), greatest (2), basic (2), win (2), shadow (2), often (2), line (2), charter (2), disabled (2), yet (2), end (2), copies (2), fills (2), our (2), thus (2), everything (2), because (2), offload (2), other (2), tasks (2), entirely (2), rasterisation (2), client (2), little (2), case (2), virtually (2), bandwidth (2), dynamic (2), i830 (2), i845 (2), garbage (2), stack (2), first (2), solution (2), reserved (2), usage (2), stable (2), every (2), before (2), extra (2), after (2), firefox (2), held (2), worse (2), key (2), gt1 (2), hd5770 (2), generations (2), q8440 (2), q35 (2), device (2), netbook (2), situation (2), complex (2), when (2), why (2), thread (2), limited (2), caches (2), ipc (2), throttled (2), cooling (2), perhaps (2), 35w (2), previous (2), their (2), brethen (2), put (2), chip (2), 6th (2), quad (2), fixed (2), beast (2), slightly (2), sporting (2), everytime (2), around (2), paintball (2), form (2), nouveau (2), several (2), ion (2), relative (2), merits (2), thanks (2), turbo (2), regression (2), gain (2), adding (2), mind (2), show (2), applications (2), feed (2), application (2), multiple (2), renderer (2), argues (2), utilize (2), available (2), work (2), enough (2), inside (2), onto (2), performs (2), default (2), blt (2), 4950hq (2), particular (2), started, name, email, comment, loading, collapse, manage, subscriptions, reader, report, content, privacy, already, free, blog, meta, rss, feeds, org, blogroll, search, 2007, september, march, april, archives, uncategorized, pages, averages, micro, extremely, 284, 128, 108, 117, looking, tests, rough, feel, xorg, 965gm, thinkpad, t61, discarding, anything, left, light, shell, investigating, heyday, ran, modern, had, changed, passed, gradually, parties, remains, recover, threw, away, performed, static, carved, scanout, important, renderbuffers, dri, clients, upside, scheme, meant, locations, tell, therefore, predetermined, resized, second, monitor, needed, framebuffer, game, wanted, relieved, role, task, guise, userspace, patch, insert, correct, addresses, relocation, bottleneck, outset, complaining, loss, retro, point, mere, midlayer, whose, existence, hinder, doing, right, thing, implementation, devil, details, fallbacks, don, huge, hindsight, decision, flawed, mappable, unmappable, cannot, map, exist, argument, fallacy, advent, tends, mapping, care, amoritize, mention, asymmetry, upload, download, speeds, exploit, transfers, strategy, regard, originally, removed, argued, unified, migrate, allocations, necessary, mapped, uncached, readback, outstanding, regularly, asked, differences, short, catastrophic, regressed, failed, forward, until, maturity, consistency, round, looked, ums, turn, arguments, reason, behind, dropping, support, writing, address, advanced, live, help, standard, graph, absolute, trace, shorter, bars, quicker, pretty, week, saw, display, translating, feeding, them, packs, standout, features, trapezoid, shaders, handle, supported, development, emphasis, measured, latest, appears, moving, windows, marginal, buffer, times, fails, seems, bring, nothing, misery, 277000, 265000, rgb, 312000, 6740, 382, 268000, 7260, 376, 154000, 1880, 308000, 6500, 380, xvfb, eye, basics, uploads, know, achieving, goals, premise, choice, complicated, consider, efficacy, efficiency, happier, result, avoided, achieve, parity, worst, less, excel, recent, cpus, process, designed, sad, fact, inadequate, challenge, workarounds, place, suites, brought, next, way, enable, eventually, eating, raison, etre, critical, requirement, cunning, reuse, stopped, streamer, seeing, remained, hours, thrashing, daniel, vetter, extended, implement, workaround, whereby, copied, area, compromised, avoid, assume, responsibility, ensuring, coherent, intervene, non, cooperative, i915, broke, hearts, fear, decade, capable, 845g, rises, phoenix, ashes, except, xserver, working, git, 20121228, methods, latency, metric, note, currently, performant, trying, shown, makes, graphs, difficult, numbers, confusingly, called, hd2000, discrete, doubling, greater, across, continue, disservice, ourselves, gen3, pineview, hear, cry, will, avx2, slight, variations, executing, wide, variation, feature, instruction, sets, necessity, different, oddity, underperforming, badly, base, clock, tiny, acheive, theory, coupled, guess, poor, glitch, broken, configuration, underrated, cetera, success, story, beasts, frankly, impressive, course, dwarfed, hands, 3770k, runs, 7ghz, 2ghz, 6ghz, cooler, 45w, respectively, gt2, variants, 7th, 66ghz, frequency, venerable, 3rd, precise, beaten, 83ghz, 4th, g45, q8400, implementations, unexpected, tale, nearly, html5, demo, fix, resource, starvation, giving, order, magnitude, strangely, fares, substantially, consistently, whereas, experience, severe, slowdowns, mixed, released, cater, low, market, substantial, upgrade, anemic, gma950, 945gm, shipped, atoms, nice, machines, runtimes, touch, side, full, grown, laptops, comparable, premium, tablets, later, occasional, request, releases, bound, example, fishbowl, fishtank, particles, eliminated, notable, beats, let, merely, tuned, resources, say, half, regress, third, bad, regressions, unacceptable, offer, tantalising, suggestion, significant, minor, remember, others, tried, ultrabook, soon, active, entire, package, thermal, constraints, dire, benchmarking, prototype, model, threading, øyvind, kolås, experiment, promise, raw, throwing, beat, simd, compositing, routines, provided, raise, whether, impacting, immediate, mode, nature, alteration, preserve, semantics, break, individual, composite, scan, conversion, pieces, pool, wait, returning, risk, overhead, outweighs, splitting, vector, søren, sandmann, pedersen, maintainer, heart, ideally, either, rate, limiting, component, attached, agent, along, might, whilst, opportunity, bear, considerably, fraction, fully, variety, common, bilinear, scaling, compute, fail, furthermore, transformation, scale, translation, operates, destination, must, incur, ignore, making, further, optimisations, improve, unit, happens, throw, zoidberg, summary, offers, meagre, attainable, takes, match, inherent, inefficiencies, succeeds, offloading, letting, uses, disable, allow, engine, data, none, multithreaded, specialised, backends, normalized, whole, jump, mostly, above, beyond, expected, likely, higher, thermals, fold, context, laptop, spiel, toy, processsor, special, iris, pro, 5200, gt3e, units, 128mib, edram, serve, fourth, glamorous, refresh, based, taken, synthetic, significantly, metrics, positive, spin, closely, clear, benchmark, crashing, noaccel, comparison, chosen, checking, progress, state, play, sitting, red, corner, skip, navigation,
Text of the page (random words):
without alteration to preserve the existing semantics we can break up the individual composite and scan conversion operations into small pieces and feed those to a pool of threads and then wait for the threads to complete before returning back to the application as such we then never run the threads for very long and risk that the overhead in thread management outweighs any benefit from splitting the operation over multiple cores to gain the greatest advantage from adding threads i used a desktop sandybridge processor an i5 2500 for benchmarking this prototype cairo image backend about half the test cases regress by about 10 but around one third are improved by 2 3x not bad the regressions are unacceptable but it does offer a tantalising suggestion that we can make some significant improvements with only a minor change remember not all processors are equal and thanks to sandybridge s turbo some cores are more equal than others i tried running the test cases on a sandybridge ultrabook and as soon as more one core was active the entire package was throttled to keep it within its thermal constraints the performance regression was dire having tuned the software backend to make better use of the average resources what can we now say about the relative merits of the integrated graphics processor indeed for the cases that are almost entirely gpu bound for example the firefox fishbowl fishtank paintball particles we have virtually eliminated all the previous advantage that the gpu held in a notable couple of cases we have improved the image backend to outperform sna and for all cases now the threaded image backend beats uxa however as can be seen there is still plenty of room for improvement of the image backend and we can t let the hardware acceleration be merely equal to a software rasteriser any old ion january 22 2013 2 13 pm posted in cairo comments 2 the nvidia ion gpu was released a few years back to cater for the low power netbook market as a substantial upgrade for the anemic intel gma950 or 945gm as it is better known to us that shipped as the integrated graphics processor in the first atoms those made quite nice machines with long runtimes a touch on the slow side compared to the full grown laptops but actually quite comparable to the current generation of premium tablets many years later people still use these and i get the occasional request to see how well they perform everytime nvidia releases a new driver so lets take a look in nearly every test there is a small improvement at around 2 3 but for the paintball html5 demo they have managed to fix some form of resource starvation and have been able to accelerate it giving an order of magnitude speed increase strangely the nvidia driver fares worse on average than the nouveau driver despite it having substantially better performance in several cases this is because the nouveau driver is at least consistently slow whereas the nvidia driver also experience a few severe slowdowns and so has a much more mixed set of results a tale of 3 processors january 4 2013 3 05 pm posted in cairo comments 3 as is always the case everytime you try to compare implementations you find completely unexpected bugs the idea was simple look at the performance of the last few generations and see how we ve been improving to start with i have a core2 quad q8400 at 2 66ghz fixed frequency and a venerable 3rd generation intel gpu a q35 to be precise this is a 95w desktop beast only to be beaten by a slightly older q9550 core2 quad at 2 83ghz sporting a 4th generation gpu a g45 from the current generation of cpu design i have two mobile chips a sandybridge i5 2520m at up to 3 2ghz and an ivybridge i7 3720qm at up to 3 6ghz these mobile chips run much cooler than their desktop brethen at 35w and 45w respectively and both have the gt2 variants of the 6th and 7th generation intel gpus to put those two into perspective i have also added the results from a desktop sandybridge chip the 95w i5 2500 which runs up to 3 7ghz on the downside this processor only has a gt1 6th generation gpu compared to the baseline performance given by the core2 q8440 q9550 1 6x faster i5 2500 2 3x faster i5 2520m 1 7x faster i7 3720qm 1 4x faster the success story is that the 35w mobile processors really do deliver the performance of the 95w desktop beasts of the previous generation which is quite frankly very impressive of course they are still dwarfed by their desktop brethen now i really do want to get my hands on a i7 3770k the oddity is then why is my ivybridge underperforming its single thread performance which is under test here should not be as badly limited the base clock is a tiny bit faster than the sandybridge i5 and with the larger and improved caches it should acheive better ipc as well in theory it should also be coupled to faster main memory as well my guess is that it is being throttled due to poor cooling or that there is a glitch in the software or perhaps a broken configuration or perhaps the memory is underrated et cetera but what of the graphics performance i hear you cry the situation here is a bit more complex all of those processors where using the same sse2 backend in pixman which will only be improved upon with the introduction of avx2 in haswell and so we were directly comparing slight variations of processor design executing the same software when we look at gpu performance not only do we have a wide variation of processor design feature set and instruction sets we also by necessity have different software for each if we look at the current driver situation that is using uxa compared to the baseline performance given by using sna on the core2 q8440 with a q35 a gen3 device like found in the pineview netbook gpu q9550 4 3x slower i5 2500 1 5x slower i5 2520m 3 0x slower i7 3720qm 2 3x slower despite almost a doubling of cpu power and an even greater increase in the gpu performance across the generations the drivers continue to do a disservice to the hardware and ourselves sandybridge acceleration december 30 2012 11 32 am posted in 2d cairo comments 6 lots of numbers from running cairo traces on a i5 2500 which has a sandybridge gt1 gpu or more confusingly called hd2000 and also a radeon hd5770 discrete gpu the key results being the geometric mean of all the traces as compared to using a software rasteriser uxa snb 1 6x slower glamor snb 2 7x slower sna snb 1 6x faster gl snb 2 9x slower exa r600g 3 8x slower glamor r600g 3 0x slower gl r600g 2 7x slower fglrx xlib 4 7x slower fglrx gl 32 0x slower not shown as it makes the graphs even more difficult to read all bar one of the acceleration methods is worse performance power latency by any metric than simply using the cpu and rendering directly within the client note also that software rasterisation is currently more performant than trying to use the gpu through the opengl driver stack all software except for the xserver which was held back to keep glamor working from git as of 20121228 845g rises like a phoenix from the ashes of xaa december 17 2012 11 55 am posted in 2d comments 8 the introduction of kms and gem into the i915 driver broke the i830 i845 chipsets and a lots of hearts but fear not a decade after its introduction we finally have a driver that is not only stable but capable of accelerating firefox the problem the problem was simply we could not find a way to enable dynamic video memory on the ancient i830 i845 chipsets without it eventually eating garbage since dynamic memory management was the raison d etre of gem and critical for acceleration it is a requirement of the current driver stack the first cunning solution was simply never to reuse batch buffers and keep a small amount of memory reserved for our usage this stopped the command streamer from seeing the garbage and my system has remained stable for many hours of thrashing daniel vetter extended my solution to implement a kernel workaround whereby every batch would be copied into a reserved area before execution in the end we compromised so that i could avoid that extra copy and assume responsibility in the driver for ensuring the batch was coherent but the kernel would intervene for any non cooperative driver with these workarounds in place we are finally able to run through the test suites which brought us to the next problem the sad fact is that uxa is inadequate for the challenge of accelerating the render protocol if we compare with an architecture that was designed to accelerate cairo sna we find a much happier result in all cases the performance is at least as good as using a software rasteriser in the x server and often much better than if we avoided the render protocol entirely and did the rasterisation in the client with a little more tuning we may be able to achieve parity even in the worst case if we can win on an old gpu with an ancient cpu single core virtually no cache and even less memory bandwidth we should be able to excel on more recent gpus and cpus and be more efficient in the process yet the render protocol is not the be all and end all of acceleration we need to keep an eye on the basics as well the copies the fills and the uploads to know if we are achieving our goals the basic premise is that using the driver and thus the gpu is faster than just using the cpu for everything in reality the choice is more complicated because we have to consider the efficacy of gpu offload for enabling the cpu to get on with other tasks and overall power efficiency 1 baseline performance of xvfb 2 sna with acceleration disabled shadow 3 uxa 4 sna 1 2 3 4 operation 277000 0 2 04 1 06 4 58 char in 80 char aa line charter 10 265000 0 2 15 1 11 4 83 char in 80 char rgb line charter 10 312000 0 0 66 0 15 1 38 copy 10x10 from window to window 6740 0 0 90 1 56 1 75 copy 100x100 from window to window 382 0 0 92 1 30 1 36 copy 500x500 from window to window 268000 0 0 74 0 17 1 50 copy 10x10 from window to pixmap 7260 0 0 87 1 43 1 85 copy 100x100 from window to pixmap 376 0 0 94 1 28 1 37 copy 500x500 from window to pixmap 154000 0 0 74 0 69 0 86 putimage 10x10 square 1880 0 1 04 1 05 1 04 putimage 100x100 square 87 1 1 03 1 02 1 01 putimage 500x500 square 308000 0 0 58 0 46 0 66 shmputimage 10x10 square 6500 0 1 02 1 16 1 24 shmputimage 100x100 square 380 0 1 00 1 24 1 28 shmputimage 500x500 square so it appears that using the gpu for basic operations such as moving the windows about is only at most a marginal win over using a shadow buffer and often times uxa fails at even that overall then it seems that enabling uxa should bring nothing but misery glamor 0 5 august 17 2012 2 08 pm posted in 2d comments 6 last week saw a new release of glamor a library for use by the 2d display drivers translating render drawing commands into opengl and then feeding them into a 3d driver this release packs in a couple of standout features for cairo trapezoid shaders and the ability to handle textures larger than supported by hardware the development emphasis has been on performance and indeed glamor 0 5 is much improved over glamor 0 4 as measured on intel s latest and greatest ivybridge architecture this is a graph of absolute time for each trace shorter bars are quicker and as can be seen the improvement is pretty good however to keep this in perspective lets compare against the standard 2d driver still plenty of room for improvement in the beginning august 1 2012 10 13 pm posted in 2d comments 16 having looked at the impact of the move from xaa to ums exa and then to kms uxa on performance of the core drawing operations we can turn to look at the impact upon render acceleration one of the arguments for the reason behind dropping xaa support and writing a new architecture was to address the needs of accelerating the new advanced rendering protocol render did the claim really live up to reality did the switch to exa actually help in short the switch to exa was catastrophic in the beginning the new acceleration architecture regressed performance in both the core rendering operations and failed to deliver the claim of improving render acceleration fast forward a few years and through improvements to the render protocol and many improvements to the driver and we reach uxa where we start to actually see some benefit from enabling gpu acceleration but it is not until we look at sna that we reach a level of maturity and consistency in the driver to have all round performance that is finally at least as good as xaa again effectively software rendering in these benchmarks an outstanding question that is regularly asked is what are design differences between exa uxa and sna uxa was originally exa with the pixmap migration removed it was argued that given a unified memory architecture such as found on igp there was no need to migrate pixmaps between the various gpu memory domains instead all pixmap allocations were to be made in the single gpu domain and if necessary the pixmap would be mapped and read by the cpu through the gtt effectively an uncached readback very very slow in hindsight that decision was flawed as it turns out not only we have both mappable and unmappable memory domains within the igp and so we cannot simply map any gpu pixmap without cost but we can also have snoopable gpu memory memory that can exist in the cpu cache the single gpu memory domain argument was a fallacy from the start and even more so with the advent of the shared last level cache between the cpu and gpu also it tends to be much much faster to copy from the gpu pixmap into system memory perform the operation on the copy in system memory and then copy it back than it is to try and perform the operation through a gtt mapping of a gpu pixmap more so if you take care to amoritize the cost of the migration not to mention the asymmetry between upload and download speeds and that we can exploit snoopable gpu memory to accelerate those transfers so it turns out that having a very efficient pixmap migration strategy is the core of an acceleration architecture in this regard sna is very much like exa from the design point of view the only real difference between exa and sna is that exa is a mere midlayer whose existence is to hinder the driver from doing the right thing and sna is a complete implementation since the devil is in the details fallbacks are slow don t that difference turns out to be quite huge going retro july 23 2012 8 50 am posted in 2d comments 3 a few years ago memory management for the graphics card was performed as single static allocation in the x server which was then carved up into surfaces and used for the scanout and important pixmaps such as renderbuffers for its dri clients the upside of this simple scheme meant that all locations were known allocation was very fast and we could always tell the gpu where its surfaces were the downside was that the amount of video memory was therefore predetermined and could not be resized not even if you added a second monitor and needed a new framebuffer or...
|