Meta tags:
author= Sicheng Zhu;
Headings (most frequently used words):
-
Text of the page (most frequently used words):
sicheng (14), zhu (14), arxiv (11), furong (11), huang (11), bang (10), code (8), and (8), proceedings (6), the (6), from (6), 2024 (6), robustness (5), models (5), icml (4), publications (4), adversarial (4), data (4), michael (4), andrei (4), panaitescu (4), liess (4), yuancheng (4), large (4), language (4), generation (4), website (3), zhang (3), model (3), 2025 (3), for (3), university (3), research (3), david (2), evans (2), learning (2), robust (2), neurips (2), generalization (2), out (2), distribution (2), iclr (2), souradip (2), chakraborty (2), text (2), aakriti (2), agrawal (2), tom (2), goldstein (2), zora (2), che (2), pankayaraj (2), pathmanathan (2), attacks (2), copyright (2), ruiyi (2), colm (2), reward (2), with (2), llm (2), phd (2), was (2), scholar (2), prof (2), received (2), template, credit, xiao, 2020, adversarially, representations, via, worst, case, mutual, information, maximization, 2021, understanding, benefit, invariance, perspective, sanghyun, hong, 2023, unforeseen, using, equivariant, domain, translator, chaithanya, kumar, mummadi, more, context, less, distraction, visual, classification, inferring, conditioning, contextual, attributes, amrit, singh, bedi, dinesh, manocha, possibilities, generated, detection, mucong, ding, tahseen, rabbani, chenghao, deng, abdirisak, mohamed, yuxin, wen, benchmarking, image, watermarks, aaai, best, paper, award, advml, frontiers, workshop, can, watermarking, prevent, copyrighted, hide, training, yigitcan, kaya, openreview, naacl, poisonedparrot, subtle, poisoning, elicit, infringing, content, gang, joe, barrow, zichao, wang, ani, nenkova, tong, sun, media, coverage, unofficial, autodan, interpretable, gradient, based, dataset, automatic, pseudo, harmful, prompt, evaluating, false, refusals, udari, madhushani, sehwag, alec, koppel, sumitra, ganesh, genarm, guided, autoregressive, test, time, alignment, brandon, amos, yuandong, tian, chuan, guo, ivan, evtimov, 2412, 10321, advprefix, objective, nuanced, jailbreaks, safety, controllable, before, visiting, working, electronic, science, technology, china, institute, electronics, chinese, academy, sciences, virginia, where, advised, interned, meta, genai, fair, adobe, bosch, maryland, college, park, bio, focus, making, settings, ultimate, goal, achieve, this, baking, symmetries, like, equivariance, into, architectures, building, design, interest, member, technical, staff, openai, team, twitter, google, sczhu, umd, edu,
Text of the page (random words):
sicheng zhu sicheng zhu sczhu umd edu google scholar twitter i am a member of technical staff at openai on the adversarial robustness research team research interest i focus on making ai models robust in out of distribution and adversarial settings my ultimate goal is to achieve this by baking symmetries like equivariance into model architectures i e building robustness by design bio i received my phd in cs from the university of maryland college park where i was advised by prof furong huang i ve interned at meta genai and fair adobe research and bosch ai before my phd i was a visiting scholar at the university of virginia working with prof david evans i received my m e from institute of electronics chinese academy of sciences and b s from university of electronic science and technology of china publications llm safety and controllable generation advprefix an objective for nuanced llm jailbreaks sicheng zhu brandon amos yuandong tian chuan guo ivan evtimov arxiv 2412 10321 arxiv code genarm reward guided generation with autoregressive reward model for test time alignment yuancheng xu udari madhushani sehwag alec koppel sicheng zhu bang an furong huang sumitra ganesh iclr 2025 proceedings arxiv automatic pseudo harmful prompt generation for evaluating false refusals in large language models bang an sicheng zhu ruiyi zhang michael andrei panaitescu liess yuancheng xu furong huang colm 2024 proceedings arxiv website code dataset autodan interpretable gradient based adversarial attacks on large language models sicheng zhu ruiyi zhang bang an gang wu joe barrow zichao wang furong huang ani nenkova tong sun colm 2024 proceedings arxiv unofficial code media coverage publications copyright poisonedparrot subtle data poisoning attacks to elicit copyright infringing content from large language models michael andrei panaitescu liess pankayaraj pathmanathan yigitcan kaya zora che bang an sicheng zhu aakriti agrawal furong huang naacl 2025 openreview can watermarking large language models prevent copyrighted text generation and hide training data michael andrei panaitescu liess zora che bang an yuancheng xu pankayaraj pathmanathan souradip chakraborty sicheng zhu tom goldstein furong huang neurips 2024 advml frontiers workshop best paper award aaai 2025 arxiv benchmarking the robustness of image watermarks bang an mucong ding tahseen rabbani aakriti agrawal yuancheng xu chenghao deng sicheng zhu abdirisak mohamed yuxin wen tom goldstein furong huang icml 2024 arxiv website code on the possibilities of ai generated text detection souradip chakraborty amrit singh bedi sicheng zhu bang an dinesh manocha furong huang icml 2024 arxiv publications generalization more context less distraction visual classification by inferring and conditioning on contextual attributes bang an sicheng zhu michael andrei panaitescu liess chaithanya kumar mummadi furong huang iclr 2024 arxiv code learning unforeseen robustness from out of distribution data using equivariant domain translator sicheng zhu bang an furong huang sanghyun hong icml 2023 proceedings code understanding the generalization benefit of model invariance from a data perspective sicheng zhu bang an furong huang neurips 2021 proceedings arxiv code publications adversarial robustness learning adversarially robust representations via worst case mutual information maximization sicheng zhu xiao zhang david evans icml 2020 proceedings arxiv code website template credit
|