Meta tags:
description= Apache Spark is a multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.;
Headings (most frequently used words):
unified, engine, for, large, scale, data, analytics, what, is, apache, spark,
Text of the page (most frequently used words):
spark (44), the (23), and (19), data (18), apache (16), sql (14), json (11), run (10), for (9), scale (7), select (7), docker (7), #engine (6), machine (6), learning (6), name (6), age (6), trademarks (5), use (5), with (5), analytics (5), first (5), logs (5), read (5), opt (5), bin (5), now (5), show (5), project (4), other (4), features (4), science (4), where (4), scala (4), connect (4), software (3), foundation (3), source (3), community (3), from (3), execution (3), large (3), your (3), thousands (3), machines (3), filter (3), model (3), pyspark (3), streaming (3), unified (3), logo (2), registered (2), under (2), license (2), committers (2), issue (2), code (2), mailing (2), open (2), contributors (2), documentation (2), join (2), tpc (2), queries (2), without (2), adaptive (2), query (2), structured (2), unstructured (2), such (2), same (2), ansi (2), algorithms (2), distributed (2), most (2), scalable (2), dataset (2), shell (2), summary (2), filtered_df (2), generate (2), accountbalance (2), csv (2), test_df (2), test (2), train_df (2), fit (2), train (2), label (2), pip (2), install (2), java (2), python (2), clusters (2), fast (2), language (2), release (2), policy (2), resources (2), feather, are, either, united, states, countries, see, guidance, all, marks, mentioned, may, their, respective, owners, copyright, 2018, licensed, version, tracking, how, contribute, news, events, list, has, thriving, around, globe, building, assisting, users, accelerates, 1tb, stats, works, tables, images, you, already, comfortable, support, adapts, plan, runtime, automatically, setting, number, reducers, built, advanced, hood, storage, infrastructure, integrates, favorite, frameworks, helping, them, ecosystem, companies, including, fortune, 500, over, 000, industry, academia, widely, used, computing, head, path, sparkr, val, last_name, last, first_name, statistics, countofdependents, subset, balance, true, header, accounts, transform, predictions, training, 100, numtrees, randomforestregressor, set, hyperparameters, algorithm, seed, randomsplit, split, into, datasets, createdataframe, every, record, contains, feature, vector, quickstart, python3, official, image, laptop, fault, tolerant, perform, exploratory, analysis, eda, petabyte, having, resort, downsampling, execute, dashboarding, hoc, reporting, runs, faster, than, warehouses, unify, processing, batches, real, time, using, preferred, batch, key, simple, multi, executing, engineering, single, node, what, get, started, event, thanks, sponsorship, homepage, website, kubernetes, operator, swift, rust, github, security, process, versioning, useful, developer, tools, developers, privacy, history, powered, tracker, improvement, proposals, spip, contributing, lists, examples, frequently, asked, questions, older, versions, latest, third, party, projects, graphx, deprecated, mllib, pandas, dataframes, libraries, download,
Text of the page (random words):
apache spark unified engine for large scale data analytics download libraries sql and dataframes spark connect spark streaming pandas on spark mllib machine learning graphx deprecated third party projects documentation latest release older versions and other resources frequently asked questions examples community mailing lists resources contributing to spark improvement proposals spip issue tracker powered by project committers project history privacy policy developers useful developer tools versioning policy release process security github spark spark connect go spark connect rust spark connect swift spark docker spark kubernetes operator spark website apache software foundation apache homepage license sponsorship thanks event unified engine for large scale data analytics get started what is apache spark apache spark is a multi language engine for executing data engineering data science and machine learning on single node machines or clusters simple fast scalable unified key features batch streaming data unify the processing of your data in batches and real time streaming using your preferred language python sql scala java or r sql analytics execute fast distributed ansi sql queries for dashboarding and ad hoc reporting runs faster than most data warehouses data science at scale perform exploratory data analysis eda on petabyte scale data without having to resort to downsampling machine learning train machine learning algorithms on a laptop and use the same code to scale to fault tolerant clusters of thousands of machines python sql scala java r run now install with pip pip install pyspark pyspark use the official docker image docker run it rm spark python3 opt spark bin pyspark quickstart machine learning analytics data science df spark read json logs json df where age 21 select name first show every record contains a label and feature vector df spark createdataframe data label features split the data into train test datasets train_df test_df df randomsplit 80 20 seed 42 set hyperparameters for the algorithm rf randomforestregressor numtrees 100 fit the model to the training data model rf fit train_df generate predictions on the test dataset model transform test_df show df spark read csv accounts csv header true select subset of features and filter for balance 0 filtered_df df select accountbalance countofdependents filter accountbalance 0 generate summary statistics filtered_df summary show run now docker run it rm spark opt spark bin spark sql spark sql select name first as first_name name last as last_name age from json logs json where age 21 run now docker run it rm spark opt spark bin spark shell scala val df spark read json logs json df where age 21 select name first show run now docker run it rm spark opt spark bin spark shell scala dataset df spark read json logs json df where age 21 select name first show run now docker run it rm spark r opt spark bin sparkr df read json path logs json df filter df df age 21 head select df df name first the most widely used engine for scalable computing thousands of companies including 80 of the fortune 500 use apache spark over 2 000 contributors to the open source project from industry and academia ecosystem apache spark integrates with your favorite frameworks helping to scale them to thousands of machines data science and machine learning sql analytics and bi storage and infrastructure spark sql engine under the hood apache spark is built on an advanced distributed sql engine for large scale data adaptive query execution spark sql adapts the execution plan at runtime such as automatically setting the number of reducers and join algorithms support for ansi sql use the same sql you re already comfortable with structured and unstructured data spark sql works on structured tables and unstructured data such as json or images tpc ds 1tb no stats with vs without adaptive query execution accelerates tpc ds queries up to 8x join the community spark has a thriving open source community with contributors from around the globe building features documentation and assisting other users mailing list source code news and events how to contribute issue tracking committers apache spark spark apache the apache feather logo and the apache spark project logo are either registered trademarks or trademarks of the apache software foundation in the united states and other countries see guidance on use of apache spark trademarks all other marks mentioned may be trademarks or registered trademarks of their respective owners copyright 2018 the apache software foundation licensed under the apache license version 2 0
|