Meta tags:
description= GitHub is where people build software. More than 150 million people use GitHub to discover, fork, and contribute to over 420 million projects.;
Headings (most frequently used words):
this, awesome, repositories, saved, searches, data, augmentation, navigation, to, your, topic, footer, snorkel, nvidia, webdataset, torchio, iver56, audiomentations, search, code, users, issues, pull, requests, provide, feedback, menu, use, filter, results, more, quickly, here, are, 582, public, matching, improve, page, add, repo, team, dali, zhaoj9014, face, evolve, qdata, textattack, project, nemo, datadesigner, 425776024, nlpcda, visual, layer, fastdup, jasonwei20, eda_nlp, agamiko, review, yongzhuo, nlp_xiaojiang, lirongwu, graph, self, supervised, learning, zhanlaoban, eda_nlp_for_chinese, tebmer, knowledge, distillation, of, llms, torch, paperspace, dataaugmentationforobjectdetection, quqxui, llm4ie, papers, goru001, inltk,
Text of the page (most frequently used words):
data (57), #augmentation (47), learning (37), code (29), python (24), updated (22), #issues (22), pull (21), requests (21), star (21), for (17), and (15), image (15), github (13), deep (13), audio (11), machine (11), nlp (10), processing (10), all (10), security (9), 2026 (9), your (9), pytorch (9), detection (9), the (8), language (8), training (7), text (7), you (6), this (6), with (6), topic (6), support (6), distillation (6), chinese (6), classification (6), search (6), face (6), enterprise (6), view (6), topics (5), sentence (5), large (5), knowledge (5), extraction (5), supervised (5), self (5), explore (5), adversarial (5), community (4), more (4), 2024 (4), models (4), graph (4), awesome (4), papers (4), apr (4), 2025 (4), sound (4), mar (4), feedback (4), eda (4), from (4), quality (4), use (4), discussions (4), medical (4), webdataset (4), stars (4), can (3), that (3), not (3), information (3), manage (3), navigation (3), page (3), add (3), embeddings (3), similarity (3), developer (3), event (3), recognition (3), generative (3), llms (3), object (3), audiomentations (3), useful (3), model (3), survey (3), neural (3), transfer (3), enhance (3), network (3), review (3), resources (3), nlpcda (3), jul (3), nemo (3), nvidia (3), high (3), library (3), textattack (3), gpu (3), snorkel (3), most (3), repositories (3), another (3), tab (3), window (3), refresh (3), session (3), reload (3), sign (3), saved (3), documentation (3), grade (3), features (3), copilot (3), platform (3), solutions (3), footer (2), learn (2), repository (2), repo (2), links (2), developers (2), about (2), indic (2), languages (2), natural (2), provide (2), out (2), box (2), application (2), nov (2), shot (2), jupyter (2), notebook (2), effects (2), dsp (2), music (2), fast (2), iver56 (2), llm (2), synthesis (2), alignment (2), paper (2), nlp数据增强 (2), aug (2), xlnet (2), augment (2), bert (2), feature (2), 相似度 (2), generation (2), find (2), here (2), visualization (2), tools (2), visual (2), analysis (2), fastdup (2), tool (2), generate (2), insights (2), synthetic (2), mcp (2), work (2), lab (2), applications (2), torchio (2), performance (2), system (2), small (2), attacks (2), weak (2), supervision (2), quickly (2), recently (2), fewest (2), forks (2), sort (2), 582 (2), filter (2), sponsors (2), events (2), collections (2), trending (2), signed (2), appearance (2), settings (2), cancel (2), see (2), available (2), searches (2), business (2), advanced (2), open (2), source (2), customer (2), services (2), devops (2), app (2), merge (2), perform, action, time, share, personal, cookies, contact, docs, status, privacy, terms, inc, associate, visit, landing, select, curate, description, easily, improve, load, jan, encoding, word, toolkit, aims, various, tasks, might, need, 838, inltk, goru001, context, cross, domain, arguments, construction, few, zero, relation, named, entity, using, llm4ie, quqxui, 2020, imagine, bounding, opencv, dataaugmentationforobjectdetection, paperspace, differentiable, waveform, inspired, torch, finetuning, instruction, following, multi, modal, compression, collects, break, down, into, elicitation, algorithms, skill, vertical, tebmer, may, 2022, easy, implement, corpus, 中文语料的eda数据增强工具, 论文阅读笔记, eda_nlp_for_chinese, zhanlaoban, pretext, task, pre, networks, unsupervised, representation, tkde, graphs, contrastive, predictive, lirongwu, sep, 2021, chatbot, distance, 自然语言处理, 小姜机器人, 闲聊检索式chatbot, bert句向量, xlnet句向量, embedding, 文本分类, 实体提取, ner, bilstm, crf, 数据增强, 同义句同义词生成, 句子主干提取, mainpart, 中文汉语短文本相似度, 文本特征工程, keras, http, service调用, nlp_xiaojiang, yongzhuo, policies, augmentations, autoaugment, style, list, will, some, common, techniques, libraries, repos, others, agamiko, 2023, rnn, swap, synonyms, cnn, position, presented, emnlp, 2019, eda_nlp, jasonwei20, classfication, novelty, duplicate, curation, outlier, dataset, powerful, free, designed, rapidly, valuable, video, datasets, helps, both, images, labels, while, significantly, reducing, operation, costs, unmatched, scalability, layer, 一键中文数据增强包, bert数据增强, pip, install, 425776024, agentic, multimodal, sdg, designer, scratch, seed, datadesigner, making, well, real, world, just, computing, imaging, project, feb, format, based, problems, strong, examples, framework, https, readthedocs, master, qdata, hard, negative, mining, landmark, fine, tuning, imbalanced, convolutional, nus, tencent, artificial, intelligence, computer, vision, paddlepaddle, evolve, zhaoj9014, pipeline, paddle, tensorflow, mxnet, accelerated, containing, highly, optimized, building, blocks, execution, engine, accelerate, inference, dali, jun, slicing, labeling, science, generating, team, least, options, javascript, java, typescript, matlab, html, 703, 734, are, public, matching, message, dismiss, alert, switched, accounts, resetting, focus, create, qualifiers, our, query, name, results, submit, include, email, address, contacted, read, every, piece, take, input, very, seriously, syntax, tips, clear, users, jump, pricing, premium, ons, powered, archive, program, accelerator, maintainer, programs, fund, partners, trust, center, forum, skills, ebooks, reports, webinars, stories, type, software, development, industries, government, manufacturing, financial, healthcare, industry, cases, devsecops, modernization, case, nonprofits, startups, medium, teams, enterprises, company, size, marketplace, changelog, blog, why, stop, leaks, before, they, start, secret, protection, secure, build, fix, vulnerabilities, enforce, changes, plan, track, instant, dev, environments, codespaces, automate, any, workflow, actions, workflows, integrate, external, registry, new, direct, agents, issue, write, better, creation, toggle, menu, skip, content,
Text of the page (random words):
data augmentation github topics github skip to content navigation menu toggle navigation sign in appearance settings platform ai code creation github copilot write better code with ai github copilot app direct agents from issue to merge mcp registry new integrate external tools developer workflows actions automate any workflow codespaces instant dev environments issues plan and track work code review manage code changes code quality enforce quality at merge application security github advanced security find and fix vulnerabilities code security secure your code as you build secret protection stop leaks before they start explore why github documentation blog changelog marketplace view all features solutions by company size enterprises small and medium teams startups nonprofits by use case app modernization devsecops devops ci cd view all use cases by industry healthcare financial services manufacturing government view all industries view all solutions resources explore by topic ai software development devops security view all topics explore by type customer stories events webinars ebooks reports business insights github skills support services documentation customer support community forum trust center partners view all resources open source community github sponsors fund open source developers programs security lab maintainer community accelerator github stars archive program repositories topics trending collections enterprise enterprise solutions enterprise platform ai powered developer platform available add ons github advanced security enterprise grade security features copilot for business enterprise grade ai features premium support enterprise grade 24 7 support pricing search or jump to search code repositories users issues pull requests search clear search syntax tips provide feedback we read every piece of feedback and take your input very seriously include my email address so i can be contacted cancel submit feedback saved searches use saved searches to filter your results more quickly name query to see all available qualifiers see our documentation cancel create saved search sign in sign up appearance settings resetting focus you signed in with another tab or window reload to refresh your session you signed out in another tab or window reload to refresh your session you switched accounts on another tab or window reload to refresh your session dismiss alert message explore topics trending collections events github sponsors data augmentation star here are 1 582 public repositories matching this topic language all filter by language all 1 582 python 734 jupyter notebook 703 html 15 c 10 matlab 10 typescript 9 c 6 r 5 java 3 javascript 3 sort most stars sort options most stars fewest stars most forks fewest forks recently updated least recently updated snorkel team snorkel star 6k code issues pull requests a system for quickly generating training data with weak supervision python data science machine learning ai weak supervision snorkel labeling data augmentation training data data slicing updated jun 8 2026 python nvidia dali star 5 7k code issues pull requests a gpu accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications python machine learning deep learning neural network mxnet gpu image processing pytorch gpu tensorflow data processing data augmentation audio processing paddle image augmentation fast data pipeline updated jul 21 2026 c zhaoj9014 face evolve star 3 6k code issues pull requests high performance face recognition library on paddlepaddle pytorch machine learning computer vision deep learning pytorch artificial intelligence feature extraction supervised learning face recognition face detection tencent transfer learning nus convolutional neural network data augmentation face alignment imbalanced learning model training fine tuning face landmark detection hard negative mining updated mar 20 2025 python qdata textattack star 3 5k code issues pull requests discussions textattack is a python framework for adversarial attacks data augmentation and model training in nlp https textattack readthedocs io en master nlp security machine learning natural language processing data augmentation adversarial machine learning adversarial examples adversarial attacks updated apr 17 2026 python webdataset webdataset star 3 1k code issues pull requests a high performance python based i o system for large and small deep learning problems with strong support for pytorch deep learning pytorch data augmentation webdataset webdataset format updated feb 9 2026 python torchio project torchio star 2 4k code issues pull requests discussions medical imaging processing for ai applications python machine learning deep learning pytorch medical image computing data augmentation augmentation medical image processing medical image analysis updated jul 21 2026 python iver56 audiomentations star 2 3k code issues pull requests discussions a python library for audio data augmentation useful for making audio ml models work well in the real world not just in the lab audio python music machine learning deep learning dsp sound sound processing data augmentation augmentation audio effects audio data augmentation updated apr 13 2026 python nvidia nemo datadesigner star 2 1k code issues pull requests discussions nemo data designer generate high quality synthetic data from scratch or from seed data mcp nvidia data generation nemo sdg data augmentation synthetic data multimodal tool use llm agentic ai updated jul 21 2026 python 425776024 nlpcda star 1 9k code issues pull requests 一键中文数据增强包 nlp数据增强 bert数据增强 eda pip install nlpcda nlp data augmentation chinese data augmentation nlpcda chinese eda updated mar 18 2025 python visual layer fastdup star 1 9k code issues pull requests fastdup is a powerful free tool designed to rapidly generate valuable insights from image and video datasets it helps enhance the quality of both images and labels while significantly reducing data operation costs all with unmatched scalability visualization python machine learning image deep learning image processing dataset image classification outlier detection object detection image analysis visual search data augmentation data curation visualization tools image similarity image duplicate detection novelty detection image classfication updated apr 14 2026 python jasonwei20 eda_nlp star 1 7k code issues pull requests data augmentation for nlp presented at emnlp 2019 nlp text classification position cnn embeddings synonyms swap classification rnn sentence data augmentation updated mar 19 2023 python agamiko data augmentation review star 1 6k code issues pull requests list of useful data augmentation resources you will find here some not common techniques libraries links to github repos papers and others review machine learning survey generative adversarial network style transfer data generation data augmentation image augmentation data synthesis autoaugment audio augmentation data augmentations augmentation policies nlp augmentation graph data augmentation updated aug 14 2024 yongzhuo nlp_xiaojiang star 1 5k code issues pull requests 自然语言处理 nlp 小姜机器人 闲聊检索式chatbot bert句向量 相似度 sentence similarity xlnet句向量 相似度 text xlnet embedding 文本分类 text classification 实体提取 ner bert bilstm crf 数据增强 text augment data enhance 同义句同义词生成 句子主干提取 mainpart 中文汉语短文本相似度 文本特征工程 keras http service调用 nlp text classification distance chatbot chinese feature bert data augmentation enhance text augment xlnet updated sep 23 2021 python lirongwu awesome graph self supervised learning star 1 4k code issues pull requests code for tkde paper self supervised learning on graphs contrastive generative or predictive machine learning deep learning transfer learning representation learning unsupervised learning data augmentation graph neural networks self supervised learning pre training pretext task updated aug 15 2024 zhanlaoban eda_nlp_for_chinese star 1 4k code issues pull requests an implement of the paper of eda for chinese corpus 中文语料的eda数据增强工具 nlp数据增强 论文阅读笔记 text classification eda chinese data augmentation chinese data augmentation easy data augmentation updated may 31 2022 python tebmer awesome knowledge distillation of llms star 1 3k code issues pull requests this repository collects papers for a survey on knowledge distillation of large language models we break down kd into knowledge elicitation and distillation algorithms and explore the skill vertical distillation of llms compression feedback survey alignment self training multi modal knowledge distillation data augmentation kd data synthesis self distillation instruction following llm large language model supervised finetuning updated mar 9 2025 iver56 torch audiomentations star 1 2k code issues pull requests fast audio data augmentation in pytorch inspired by audiomentations useful for deep learning audio python music machine learning deep learning dsp waveform sound pytorch sound processing data augmentation augmentation audio effects differentiable data augmentation audio data augmentation updated nov 24 2025 python paperspace dataaugmentationforobjectdetection star 1 2k code issues pull requests data augmentation for object detection opencv deep learning object detection data augmentation bounding box imagine augmentation updated apr 14 2020 jupyter notebook quqxui awesome llm4ie papers star 1 1k code issues pull requests awesome papers about generative information extraction ie using large language models llms information extraction named entity recognition event detection event extraction data augmentation relation extraction zero shot learning few shot learning knowledge graph construction event arguments cross domain learning in context learning large language models updated nov 18 2024 goru001 inltk star 838 code issues pull requests natural language toolkit for indic languages aims to provide out of the box support for various nlp tasks that an application developer might need nlp deep learning word embeddings pytorch data augmentation indic languages sentence similarity sentence embeddings sentence encoding updated jan 20 2024 python load more improve this page add a description image and links to the data augmentation topic page so that developers can more easily learn about it curate this topic add this topic to your repo to associate your repository with the data augmentation topic visit your repo s landing page and select manage topics learn more footer 2026 github inc footer navigation terms privacy security status community docs contact manage cookies do not share my personal information you can t perform that action at this time
|