Meta tags:
description= A self-hosted Jupyter-style notebook for Apache Spark with built-in S3 storage, per-user isolation, and OAuth login. Apache-2.0.;
Headings (most frequently used words):
for, your, the, one, spark, notebooks, data, everyone, can, that, in, to, notebook, user, shared, pyspark, and, scala, notebooksfor, team, built, teams, need, not, just, apache, browser, six, things, you, don, have, build, yourself, pick, isolation, level, clone, run, log, where, it, lands, full, platform, self, hosted, customized, company, questions, asks, first, without, babysitting, legally, leave, network, surface, every, bu, properly, isolated, literally, read, greedy, take, down, room, login, through, google, or, microsoft, no, new, passwords, manage, bucket, datasets, shelf, private, dead, kernels, heal, themselves, before, users, notice, restart, backend, lose, nothing, pre, installed, libraries, docker_per_user, k8s_per_user,
Text of the page (most frequently used words):
the (51), spark (28), and (24), user (20), you (19), for (16), data (15), sparklabx (14), apache (14), one (14), with (13), kernel (13), your (13), run (13), scala (12), #notebook (11), #notebooks (11), iam (10), per (10), that (10), isolation (9), container (8), pyspark (8), helm (8), can (8), everyone (8), shared (7), from (7), workspace (7), users (7), self (7), hosted (7), pro (7), stack (7), org (7), backend (6), minio (6), gets (6), every (6), layer (6), s3a (6), public (6), source (6), full (6), platform (6), security (6), chart (6), production (6), own (6), runs (6), install (6), pod (6), their (6), delta (6), csv (6), open (5), this (5), jupyter (5), oauth (5), jupyterhub (5), kernels (5), sql (5), admin (5), true (5), read (5), gpa (5), github (4), any (4), isolated (4), are (4), storage (4), enforced (4), into (4), ingestion (4), workflows (4), governance (4), access (4), what (4), demo (4), cluster (4), team (4), under (4), control (4), wiring (4), idle (4), reaper (4), teams (4), built (4), clone (4), secrets (4), docker (4), show (4), python (4), option (4), sparksession (4), scoped (3), prefix (3), before (3), app (3), else (3), query (3), engine (3), center (3), company (3), email (3), three (3), deployment (3), laptop (3), multi (3), k8s_per_user (3), docker_per_user (3), k8s (3), start (3), just (3), ship (3), over (3), bundled (3), when (3), don (3), need (3), down (3), com (3), quickstart (3), password (3), env (3), random (3), real (3), all (3), each (3), spark_2 (3), press (3), execute (3), avg (3), students (3), import (3), pain (3), docs (2), first (2), login (2), account (2), plus (2), code (2), accessdenied (2), someone (2), how (2), actually (2), core (2), new (2), going (2), orchestrated (2), medallion (2), pipelines (2), catalog (2), customized (2), repo (2), yes (2), have (2), other (2), via (2), use (2), ready (2), path (2), essentially (2), end (2), top (2), auto (2), respawn (2), build (2), work (2), faq (2), lakehouse (2), where (2), infrastructure (2), needs (2), pip (2), deploy (2), crash (2), plugin (2), only (2), diy (2), git (2), running (2), http (2), localhost (2), 3000 (2), images (2), log (2), printed (2), generates (2), fresh (2), networkpolicy (2), dev (2), single (2), pick (2), restart (2), nobody (2), themselves (2), datasets (2), collect (2), stop (2), preview (2), aws (2), hadoop (2), libraries (2), filter (2), orderby (2), agg (2), department (2), groupby (2), inferschema (2), header (2), avg_gpa (2), val (2), version (2), getorcreate (2), jars (2), packages (2), config (2), hello (2), appname (2), builder (2), cell (2), view (2), wants (2), but (2), not (2), vpc (2), leave (2), 2026, inc, provisions, policy, launched, credentials, baked, line, reading, ever, hits, design, difference, project, filters, paths, application, username, guarantees, stays, bug, fixes, stability, development, effort, watch, announcements, early, roadmap, works, eks, gke, aks, k3s, bare, metal, long, default, storageclass, ingress, controller, modes, cover, everything, tenant, sweet, spot, roughly, active, below, vanilla, probably, enough, above, wanting, pieces, adds, architectures, scale, further, haven, pressure, tested, past, yet, big, does, make, sense, see, deployed, standard, wider, layers, case, give, today, want, perfectly, valid, anyway, dies, scratch, saves, integration, questions, asks, stay, free, released, progressively, talk, about, observability, across, monitoring, discover, document, hoc, extra, bronze, silver, gold, scheduling, backfill, history, processing, land, raw, rest, lifecycle, tailored, configurable, allowlist, filesystem, configured, plain, honest, comparison, oss, otherwise, reach, went, road, lands, alternatives, https, url, starting, pulling, generated, stored, gitignored, generating, command, projects, zsh, optional, boots, licensed, fork, customize, script, strong, pulls, starts, sensible, defaults, done, second, rbac, quotas, kubernetes, resourcequota, mode, paying, requires, sock, creds, host, daemon, great, local, prod, parity, trusted, hosts, lowest, ram, quick, demos, cross, reads, possible, same, product, shapes, move, boundaries, level, mappings, state, spawn, progress, postgres, mid, execution, keep, reconnect, transparently, lose, nothing, exited, overnight, hit, crashloopbackoff, next, connect, detects, corpse, deletes, spawns, has, page, call, engineer, 7am, dead, heal, notice, drop, reference, lives, invisible, bucket, shelf, private, wired, restrict, domain, individual, unallowed, addresses, fail, row, created, yourcompany, through, google, microsoft, passwords, manage, 100, dataframe, ooms, keeps, coding, greedy, take, room, misclick, typo, returns, never, sees, request, literally, most, setups, ships, boring, bits, between, survives, monday, morning, six, things, yourself, hood, against, add, more, require, lang, library, almond, kernel_2, amazonaws, java, sdk, bundle, streaming_2, sql_2, core_2, pre, installed, false, ascending, desc, print, println, functions, help, edit, file, disconnect, clear, output, connected, 512, schema, yaml, orders, json, events_2024, parquet, experiments, models, 640, notes, queries, sample, space, files, faithful, click, browser, try, once, onboard, business, unit, trust, them, acls, leaky, surface, properly, internal, platforms, entirely, inside, prem, auth, yours, compliance, residency, rules, dpo, concerns, rule, out, saas, pii, phi, cannot, legally, network, regulated, industries, analyst, already, means, shares, bad, takes, whole, without, babysitting, people, who, together, get, started, environment, star, compare, features, why,
Text of the page (random words):
sparklabx notebook self hosted apache spark notebooks with per user isolation why demo features compare pro docs faq star apache 2 0 self hosted pyspark and scala notebooks for your team self hosted multi user every user gets their own isolated environment one helm install apache 2 0 get started view on github built for built for teams that need spark not just notebooks three teams who pick sparklabx over wiring jupyterhub spark and an iam layer together themselves data teams 5 50 people spark notebooks for everyone without the babysitting the pain everyone needs pyspark but nobody wants to install spark on a laptop a shared jupyter vm means everyone shares one kernel one bad df collect takes the whole team down every analyst gets their own kernel container on the cluster you already run open source under apache 2 0 regulated industries data that legally can t leave your network the pain compliance residency rules or dpo concerns rule out saas notebooks pii and phi cannot leave your vpc full stop runs entirely inside your vpc or on prem k8s storage kernels auth all yours internal data platforms one notebook surface for every bu properly isolated the pain each business unit wants notebooks but you can t trust them not to read each other s data app layer acls are leaky per user s3 prefix enforced at the iam layer deploy once with helm onboard via oauth try it apache spark in your browser a faithful preview of the real notebook ui click run on any cell files my space public s3a workspace users admin datasets sample csv 2 1 kb queries sql 1 4 kb notes md 640 b models experiments students csv 1 2 mb events_2024 parquet 84 mb orders json 3 4 mb schema yaml 512 b scala python connected libraries run all clear output disconnect file edit view run cell help scala python in 1 import org apache spark sql sparksession import org apache spark sql functions _ val spark sparksession builder appname hello sparklabx config spark jars packages io delta delta spark_2 12 3 0 0 getorcreate println s spark spark version from pyspark sql import sparksession spark sparksession builder appname hello sparklabx config spark jars packages io delta delta spark_2 12 3 0 0 getorcreate print f spark spark version press run to execute scala python in 2 val df spark read option header true option inferschema true csv s3a workspace public students csv df groupby department agg avg gpa as avg_gpa orderby avg_gpa desc show df spark read option header true option inferschema true csv s3a workspace public students csv df groupby department agg gpa avg orderby avg gpa ascending false show press run to execute scala python in 3 df filter gpa 3 5 show 5 df filter gpa 3 5 show 5 press run to execute pre installed libraries org apache spark spark core_2 12 3 5 1 org apache spark spark sql_2 12 3 5 1 org apache spark spark streaming_2 12 3 5 1 io delta delta spark_2 12 3 0 0 org apache hadoop hadoop aws 3 3 4 com amazonaws aws java sdk bundle 1 12 x sh almond scala kernel_2 12 0 14 0 org scala lang scala library 2 12 18 add more with require scala or pip install pyspark this is a ui preview the production notebook runs against a real spark cluster you control under the hood six things you don t have to build yourself most self hosted notebook setups stop at docker run jupyter sparklabx ships the boring infrastructure the bits between it runs on my laptop and it survives monday morning user a literally can t read user b s data each user gets a minio iam account scoped to their own s3 prefix a misclick or a typo in a notebook path returns accessdenied from the storage layer your app code never sees the request one greedy notebook can t take down the room every user gets their own kernel container pod someone runs df collect on a 100 gb dataframe only their pod ooms everyone else keeps coding login through google or microsoft no new passwords to manage oauth wired in restrict access by domain yourcompany com or by individual email unallowed addresses fail before any db row gets created one bucket for shared datasets one shelf for everyone s private notebooks drop reference data into s3a workspace public every user can read it from pyspark or scala their own work lives at s3a workspace users you invisible to everyone else dead kernels heal themselves before users notice container exited overnight pod hit crashloopbackoff the next connect detects the corpse deletes it and spawns a fresh one nobody has to page the on call engineer at 7am restart the backend lose nothing notebook kernel mappings idle reaper state spawn progress all in postgres backend pod can crash and restart mid execution the running kernels keep going and reconnect transparently deployment pick your isolation level same product three deployment shapes start with shared for a demo move to k8s_per_user when you need real boundaries demo shared one jupyter container for everyone quick demos and single user dev cross user spark reads are possible don t ship this to production single shared kernel no isolation lowest ram trusted hosts docker_per_user one container per user on the host s docker daemon true minio iam isolation per kernel great for local dev with prod parity per user iam creds idle reaper requires docker sock access production k8s_per_user one pod per user backend runs in cluster minio iam plus kubernetes networkpolicy and resourcequota the mode you actually run for paying users pod isolation networkpolicy quotas rbac scoped backend 30 second quickstart clone run log in the script generates strong random secrets pulls public docker images and starts the stack with sensible defaults the admin password is printed when it s done 1 clone the repo apache 2 0 licensed fork customize ship 2 run quickstart sh generates a fresh env with random secrets and boots the full stack 3 open http localhost 3000 log in as admin with the printed password oauth is optional projects zsh clone git clone https github com sparklabx sparklabx git cd sparklabx one command full stack quickstart sh generating env with random secrets secrets generated stored in env gitignored pulling images starting stack backend is up sparklabx notebook is running url http localhost 3000 admin admin password kernel docker_per_user vs alternatives where it lands an honest comparison with the oss notebooks you d otherwise reach for sparklabx is essentially what you d end up wiring on top of jupyterhub if you went down that road bundled into one helm chart plain jupyter jupyterhub sparklabx pyspark scala kernels diy install diy install bundled configured multi user per user kernel container per user storage isolation filesystem only iam enforced at the s3 layer oauth email allowlist plugin plugin built in idle kernel reaper configurable built in auto respawn on kernel crash deploy pip install helm chart helm chart sparklabx pro the full data platform self hosted customized for your company notebooks are where teams start pro is the rest of the lifecycle in one workspace self hosted on infrastructure you control and tailored to your company s data stack and security needs data ingestion land raw data from any source into your lakehouse workflows processing medallion pipelines bronze silver gold with scheduling backfill and run history query engine ad hoc sql over the lakehouse no extra wiring data catalog governance discover document and control access to your data monitoring security observability and a security center across the platform talk to us about pro open source notebooks stay free under apache 2 0 platform source is released progressively faq the questions everyone asks first can t i just use jupyterhub spark yes and if you want full control over the stack that s a perfectly valid path sparklabx is essentially what you d end up wiring on top of jupyterhub anyway pyspark and scala kernels bundled per user s3 isolation enforced by minio iam an idle reaper and auto respawn when a kernel container dies if you don t need to build that stack from scratch this saves you the integration work is it production ready the open source core notebook engine kernel isolation s3 iam oauth helm chart is what you see apache 2 0 deployed via helm runs on standard k8s the wider platform layers ingestion workflows query governance security ship under pro if your use case is give my data team isolated spark notebooks you re production ready today how big a team does this make sense for sweet spot is roughly 5 to 50 active users below that vanilla jupyter on a vm is probably enough above that you ll start wanting the platform pieces pro adds ingestion orchestrated workflows governance and a security center architectures scale further we just haven t pressure tested past 50 yet can i run this on my own k8s cluster yes helm chart in chart works on eks gke aks k3s and bare metal as long as you have a default storageclass and an ingress controller three other deployment modes shared docker_per_user k8s_per_user cover everything from a laptop demo to multi tenant production what s on the roadmap the open source core stays as is bug fixes and stability the new development effort is going into pro the full self hosted data platform ingestion orchestrated workflows medallion pipelines a query engine data catalog governance and a security center customized to your company email for early access or watch the github repo for announcements how are storage isolation guarantees actually enforced on first login the backend provisions a minio iam account per user with a policy scoped to users username plus the shared public prefix the kernel container is launched with that user s credentials baked in a pyspark line reading s3a workspace users someone else gets accessdenied from minio before it ever hits any app code that s the design difference vs every isolated notebook project that filters paths in the application layer 2026 sparklabx inc docs github
|