Meta tags:
description= Simon Zirui Guo;
Headings (most frequently used words):
for, arxiv, kernels, efficient, and, research, hi, am, simon, highlights, resume, cv, kevin, multi, turn, rl, generating, cuda, kernelbench, can, llms, write, gpu, bam, just, like, that, simple, parameter, upcycling, mixture, of, experts, parallelism, in, bundle, adjustment, slam, extended, abstract, gemmini, an, open, source, full, system, dnn, accelerator, design, evaluation, platform, workshop, presentation, d3, dynamic, deadline, driven, approach, building, autonomous, vehicles, acm,
Text of the page (most frequently used words):
and (17), simon (9), for (9), guo (8), learning (8), conference (7), systems (7), international (7), computer (6), with (6), training (6), machine (6), prof (6), workshop (5), the (5), kernels (5), acm (4), research (4), efficient (4), self (3), 2022 (3), sophia (3), shao (3), arxiv (3), pre (3), icml (3), 2025 (3), currently (3), working (3), lab (3), 2026 (2), cars (2), ion (2), stoica (2), design (2), dnn (2), open (2), source (2), architecture (2), ieee (2), symposium (2), yakun (2), zirui (2), gemmini (2), slam (2), student (2), upcycling (2), moe (2), mixture (2), more (2), 2024 (2), fomo (2), zhang (2), bam (2), that (2), llms (2), gpu (2), representations (2), iclr (2), christopher (2), azalia (2), mirhoseini (2), kernelbench (2), multi (2), turn (2), generating (2), cuda (2), kevin (2), post (2), interested (2), improvement (2), models (2), scaling (2), out (2), stanford (2), time (2), drive (2), during (2), also (2), driving, robots, proceedings, european, eurosys, worked, undergraduate, assistant, author, ionel, gog, sukrit, kalra, peter, schafhalter, joseph, gonzalez, dynamic, deadline, driven, approach, building, autonomous, vehicles, presentation, space, exploration, accelerators, across, stack, first, oscar, isca, hasan, genc, seah, kim, vadim, vadimovich, nikiforov, borivoje, nikolić, krste, asanović, full, system, accelerator, evaluation, platform, extended, abstract, speed, exploiting, structural, sparsity, custom, tensor, cores, competition, microarchitecture, micro, parallelism, bundle, adjustment, attention, neural, information, processing, neurips, ngsm, spotlight, qizhen, nikolas, gritsch, dwaraknath, gnaneshwar, david, cairuz, bharat, venkitesh, jakob, foerster, phil, blunsom, sebastian, ruder, ahmet, üstün, acyr, locatelli, just, like, simple, parameter, experts, benchmark, environment, evaluate, ability, generate, dl4c, best, paper, ssi, indicates, equal, contribution, anne, ouyang, simran, arora, alex, william, can, write, exait, iii, carlo, baronio, pietro, marsella, ben, pan, silas, alberti, command, cpt, tpt, kernelllm, highlights, interplay, enables, well, improving, through, synthetic, data, published, conferences, code, generation, capability, resume, you, are, journey, please, check, rest, this, site, feel, free, contact, simonguo, dot, edu, work, centers, intersection, focus, enabling, progressing, toward, most, recently, spent, some, language, experimenting, evolutionary, algorithms, previously, designed, gpus, scaled, distributed, make, themselves, nvidia, anyscale, apple, sakana, cohere, studied, undergrad, berkeley, was, luckily, involved, rise, slice, electrical, engineering, sciences, university, leave, there, have, been, fortunate, collaborate, learn, from, tatsunori, hashimoto, ludwig, schmidt, thinking, machines, tpus, brrr, writings, publications, projects, about,
Text of the page (random words):
simon guo simon guo about projects more publications writings tpus go brrr hi i am simon i am currently at thinking machines lab i am also a cs ph d student at stanford university currently on leave working with prof azalia mirhoseini during my time there i have also been fortunate to collaborate and learn from prof ludwig schmidt prof christopher ré and prof tatsunori hashimoto i studied electrical engineering and computer sciences during my undergrad at berkeley i was luckily involved in the slice lab working with prof sophia shao and rise lab working with prof ion stoica my work centers on the intersection of computer systems and machine learning with a focus on enabling scaling and progressing toward self improvement most recently i spent some time pre training language models at cohere and experimenting with evolutionary algorithms at sakana ai previously i designed gpus at apple scaled out distributed systems at anyscale and make cars drive themselves at nvidia drive if you are interested in my journey please check out the rest of this site feel free to contact me at simonguo stanford dot edu resume cv research currently i m interested in the interplay of machine learning and systems that enables self improvement as well as improving models code generation capability through post training and scaling synthetic data i ve published at machine learning and computer systems conferences highlights ml systems kernelbench gemmini d3 post training kevin kernelllm tpt pre training cpt bam command r kevin multi turn rl for generating cuda kernels carlo baronio pietro marsella ben pan simon guo silas alberti international conference on learning representations iclr 2026 exait es fomo iii workshop at international conference on machine learning icml 2025 multi turn rl training for generating cuda kernels arxiv kernelbench can llms write efficient gpu kernels anne ouyang simon guo simran arora alex l zhang william hu christopher ré azalia mirhoseini indicates equal contribution international conference on machine learning icml 2025 dl4c best paper ssi fm workshop at international conference on learning representations iclr 2025 benchmark and environment to evaluate llms ability to generate efficient gpu kernels arxiv bam just like that simple and efficient parameter upcycling for mixture of experts qizhen zhang nikolas gritsch dwaraknath gnaneshwar simon guo david cairuz bharat venkitesh jakob foerster phil blunsom sebastian ruder ahmet üstün acyr locatelli conference on neural information processing systems neurips 2024 ngsm spotlight and es fomo ii workshop at international conference on machine learning icml 2024 upcycling moe with mixture of attention for more efficient moe pre training arxiv parallelism in bundle adjustment for slam simon zirui guo yakun sophia shao acm student research competition at ieee acm international symposium on microarchitecture micro 2022 speed up slam by exploiting structural sparsity and custom kernels on tensor cores extended abstract gemmini an open source full system dnn accelerator design and evaluation platform hasan genc seah kim vadim vadimovich nikiforov simon zirui guo borivoje nikolić krste asanović yakun sophia shao first workshop on open source computer architecture research oscar at acm ieee international symposium on computer architecture isca 2022 design space exploration for dnn accelerators across the stack workshop presentation d3 a dynamic deadline driven approach for building autonomous vehicles ionel gog sukrit kalra peter schafhalter joseph e gonzalez ion stoica worked as undergraduate research assistant for author in proceedings of european conference on computer systems eurosys 2022 os for self driving cars and robots acm 2026 simon guo
|