kernel space / systems console

I build systems that turn chaos into signal.

This is not a portfolio wall. It is a technical dossier for a builder working across data infrastructure, backend engineering, and applied AI.

M.S. Applied Data Science @ San Jose State University

Open to backend, data, and platform roles

Bias toward systems that survive real-world use

zsh — aryamehta@kernelspace
aryamehta@kernelspace ~ $ neofetch --builder-setup
OS:macOS Sequoia 15.4Host:MacBook Air 13" (Apple Silicon M4)Shell:zsh 5.9IDE:Antigravity IDE (v4.6-Thinking)Theme:Dracula DarkCursor:Active Glow Tracker (glowing: true)
aryamehta@kernelspace ~ $ cat preferences.json
{
"mode": "system-architect",
"specialization": "high-throughput-data",
"compilation": "antigravity-accelerated"
}
sys_telemetry // port_8080
signal_active

Uptime

99.987%

Latency

14ms

Ingest Rate

1.24 GB/s

live_feed_v0.9

Live build surface

graph_rag.tsretrieval_flow.yamldeploy.sh

Currently building

Agentic workflows, graph-based reasoning, and systems that feel product-grade instead of experimental.

Builder habits

Midnight coding, architecture diagrams before implementation, and making complex systems readable to both engineers and non-engineers.

What I optimize for

Clarity under pressure, strong operational defaults, and technical work that still looks good after real-world traffic, failure, and change.

Systems OS

One builder, multiple operating modes.

This is the part that should feel different. Switch modes and the system surface changes with the kind of work being done.

Mode 01

Structure chaos before the code gets expensive.

This mode is about boundaries, sequencing, and turning vague product pressure into a system shape that can actually ship.

architecture-firstproduct-grade executionclear ownership

Active case file

Enterprise RAG

Enterprise retrieval, auth, caching, and vector infrastructure delivered as one product-grade system.

Open case file

Applied AI system

Enterprise RAG

Enterprise-grade vector search, auth, and caching infrastructure.

1M+ document scale narrative
92% answer accuracy framing
LangChainPyTorchFastAPIPySpark

Identity chamber

This should feel like opening a technical mind, not scrolling a resume.

I do not want to be remembered as someone who can list tools. I want to be remembered as someone who can structure complexity.

The work I care about sits at the intersection of infrastructure, intelligence, and product clarity.

The best systems are not only fast or clever. They are legible, observable, and trusted by the people using them.

Throughput gain

0%

Refactored backend services from O(n²) to O(n log n) throughput using B+ trees and hashing on Linux services.

p95 latency drop

0%

Achieved via intelligent database indexing, query optimization, and read-replica replication schemes.

Records processed daily

0M+

Successfully handled in high-availability web and ETL platform configurations using React, FastAPI, Postgres, and Redis.

Release time cut

0%

Cut deployment overhead with automated CI/CD pipelines, blue-green releases, and zero-downtime rollouts.

Medium signals

Code is one layer. Writing is where the thinking becomes visible.

This section shows the public thought process behind the systems work, so visitors see more than repos and buzzwords.

Open Medium

Systems Engineering

Refactoring Hot Paths: How We Achieved O(n log n) Throughput using B+ Trees

A deep dive into rewriting database indexing structures and resolving query latency drag on high-throughput backend services.

JavaSpring BootB+ Trees
// Core Takeaway:Swapping linear query evaluations with structural B+ tree hashing increased throughput by 40% and cut p95 latency by 60%.

Platform Operations

Orchestrating 5 TB/Day on Kubernetes: Autoscale Tunings & Spark Partitioning Tricks

My operational checklist for scheduling distributed jobs, Celery executors, and configuring EKS node pools to optimize wall-clock efficiency.

AWSKubernetesSpark
// Core Takeaway:Tuning partitioning boundaries and predicate pushdowns decreased spark data shuffle by 60% and trimmed cloud billing costs by 40%.

Applied AI Design

Multi-Agent Travel OS: Orchestrating LangGraph State Machines & Coinbase Transactions

Connecting autonomous supervisor routing nodes to zero-hallucination vector retrievals and generative transaction payments.

PythonLangGraphFastAPI
// Core Takeaway:Binding MongoDBSaver checkpointing with CDP wallets handles persistent memory state logic outside standard LLM prompt contexts.

Product Infrastructure

Civic-Tech at Scale: Real-Time WebSockets & Dijkstra Route Optimizations

How we scaled the award-winning SJ HOPES platform to handle concurrent traffic spikes during emergency shelter allocations.

JavaSpring BootWebSocket
// Core Takeaway:Geospatial indexing (R-trees) and connection pooling cut map routing processing time by 45%.

Applied AI Design

Designing Zero-Hallucination RAG: Voyager Embeddings & MongoDB Atlas Pre-Filters

Architecting autonomous travel recovery assistants with strict term filters, context checking, and fallback search models.

RAGMongoDBFastAPI
// Core Takeaway:Pre-filtering high-dimensional vectors by exact terminal status IDs prevents LLMs from suggesting offline routes.

DevOps & SRE

Distributed Telemetry Pipelines: Proactive Monitoring with Prometheus & Grafana

Building automated logs analysis systems and setting SLO threshold indicators to decrease production incidents.

PrometheusGrafanaBash
// Core Takeaway:Automated alert rules and alert channels cut incident response times by 75% and saved 20 hours/week manual toil.

Experience log

Where the systems thinking was sharpened under real constraints.

Not a list of titles. A trace of environments where scale, reliability, and technical ownership were the daily operating conditions.

MeteoControl

Industry

Software Engineer

MeteoControl

Jan 2024 – Dec 2024Ahmedabad, IndiaJava/Spring Boot & React scale
Shipped production features weekly across Java/Spring Boot services and React/TypeScript UIs for 10k+ users, sustaining 99.9%+ availability via Prometheus/Grafana alerts and on-call rotations
Refactored hot paths from O(n²) → O(n log n) using B+ trees & hashing → +40% throughput and −60% p95 latency on Linux systems
Implemented resilience (circuit breakers, retries) and blue-green/canary deploys; automated CI/CD to cut release time 50% with zero-downtime rollouts
Built Python/Bash tooling for deploy/monitor/log analysis, eliminating ~20 hrs/week manual toil and reducing incidents 75% via proactive troubleshooting
JavaSpring BootReactTypeScriptPythonBashPrometheusGrafanaCI/CDLinux

Industry

Software Engineer Intern

Dupat Infotronicx Pvt. Ltd.

Aug 2022 – Aug 2023Ahmedabad, India1M+ records/day / 70% DB load cut
Developed an end-to-end web + data platform (React, FastAPI, Postgres/Redis) processing 1M+ records/day; caching, pagination, and rate-limiting cut DB load 70%
Re-architected ETL with memoization + batched I/O → 3× faster processing and −65% CPU
Delivered REST APIs with ~95% test coverage; created dashboards/runbooks for on-call operations
Collaborated with cross-functional teams in Agile sprints and presented technical insights to non-technical stakeholders
ReactFastAPIPostgreSQLRedisETLREST APIsAgile

Education signal

The academic bedrock underneath the systems work.

Degrees are context, not credentials. These programs shaped how I think about data, systems, and applied intelligence.

San Jose State University

Master of Science

Applied Data Science

San Jose State University

Jan 2025 – Dec 2026San Jose, CAGPA: 3.8 / 4.0

Relevant coursework

Data EngineeringMachine LearningDeep LearningData MiningBig Data TechnologiesNatural Language ProcessingData VisualizationCloud Computing

SJ Hacks 2026 Winner — civic-tech platform with 1000+ user target

MongoDB Agentic Hackathon Finalist — multi-agent travel OS

LangGraph contributor — Agentic Commerce architecture pattern

Silver Oak University

Bachelor of Technology

Computer Science & Engineering

Silver Oak University

Aug 2019 – May 2023Ahmedabad, India

Relevant coursework

Data Structures & AlgorithmsOperating SystemsDatabase ManagementComputer NetworksSoftware EngineeringObject-Oriented Programming

Led software engineering group projects

Top academic performance in CS track

Stack atlas

Not a stack list.A systems fingerprint.

Most portfolios dump tools into badges. This section maps how the technical language of the work bends toward architecture, scale, intelligence, and shipping.

surface areawide
center of gravitysystems
shipping modeproduction-minded

lane 01

Languages

01

foundation

PythonJavaGoC++TypeScriptSQL

lane 02

AI + ML

02

intelligence

🕸️LangGraph⛓️LangChain🔍RAG🤖Agentic Workflows🔥PyTorch🔶TensorFlow

lane 03

Backend + Web

03

interfaces

FastAPISpring BootNode.jsReactNext.js📡gRPC🔌WebSocketNGINX

lane 04

Data + Infra

04

data flow

PostgreSQLAWS Aurora💾DynamoDB🍃MongoDB AtlasRedis🔀KafkaSpark🌬️Airflow

lane 05

Cloud + DevOps

05

shipping

AWSDockerKubernetes🚀GitHub ActionsJenkinsPrometheusGrafanaJira

lane 06

ML Systems & Big Data

06

ml systems & scale

PySparkHadoop📈MLflow🔍Voyage AI🍃MongoDB Vector SearchScikit-learnPandasNumPy

Skills constellation

Skills are not isolated. They form a system.

Hover to explore connections. Node size reflects depth. Edges show how tools relate in real projects.

Backend
AI / ML
Data
Cloud / DevOps
Product

Operating model

Skills are not the interface. Decision patterns are.

Instead of showing a generic toolbox list, this section shows how I approach technical work from inputs to deployment.

Ingest

Map the raw signals before writing the first abstraction.

I start by understanding sources, constraints, and failure patterns. Clean systems begin with honest inputs, not optimistic assumptions.

Tools in this mode

PythonSQLKafkaREST APIsData contracts

Proof of use

Used in ETL design, data platforms, graph fraud workflows, and multi-source forecasting pipelines.

Case files

Flagship systems presented with a sharper point of view.

Each card now opens with a stronger claim. Inside, the project pages carry the architecture, tooling, impact, and delivery story.

Applied AI system

Enterprise RAG

Project

A high-performance retrieval engine for enterprise knowledge workflows.

Enterprise-grade vector search, auth, and caching infrastructure.

retrieval qualityAPI delivery

Stack + delivery

FastAPI, React, pgvector, Redis, Docker, and JWT auth.

LangChainPyTorchFastAPI
1M+ document scale narrative

Graph ML / fraud intelligence

Graph-Based Fraud Detection

Project

Graph-based intelligence pipeline exposing relational fraud rings.

Expose laundering rings and coordinated transaction abuse.

graph featuresfraud signals

Stack + delivery

Neo4j, Graph Data Science, Python, and analyst dashboard.

PythonPyTorchGraph ML
31.9M transactions modeled

Platform engineering

Distributed Data Platform

Project

ETL infrastructure built with quality rules and autoscaling discipline.

5 TB+/day ETL pipeline built for scale and operational quality.

5TB+ daily99.9% job success

Stack + delivery

Airflow, Spark, Kubernetes, Prometheus, and Grafana.

PythonAirflowSpark
5TB+ processed daily

Agentic product system

LayoverOS

Project

Multi-agent travel OS with vector search and Coinbase CDP.

Context-aware traveler recovery system with persistent memory.

agentic commercevector search

Stack + delivery

LangGraph, Atlas Vector Search, FastAPI, and Next.js.

PythonFastAPILangGraph
Agentic hackathon finalist

Data governance system

Automated Governance Catalog

Project

Automated metadata catalog for discovery and compliance.

Governance infrastructure turning tribal knowledge into system memory.

metadata automationgovernance

Stack + delivery

Python, custom catalog flows, and metadata automation.

PythonMetadata AutomationDocumentation
Metadata capture automation

Big Data ML / Trust Systems

TrustGuard AI

Project

Big Data review filter and trusted recommender engine.

6-layer multi-modal defense against bot farm manipulation.

1.98M reviews25% RMSE gain

Stack + delivery

PySpark, Spark MLlib, K-Means, ALS, and Streamlit.

PySparkSpark MLlibStreamlit
25% RMSE improvement

Machine Learning / EDA

Netflix Popularity

Project

Predictive modeling pipeline for pre-release content success.

Predict pre-release content success using synopses and metadata.

80.6% accuracy0.721 ROC-AUC

Stack + delivery

Python, scikit-learn, Random Forest, and Power BI.

Pythonscikit-learnNLTK
80% Random Forest accuracy

Deep Learning / Audio DSP

Speech Emotion Recognition

Project

Spatial-temporal deep learning speech classifier.

Detect human emotional states from spoken audio in real-time.

+15.8% accuracy gainRTF 0.00008 (Real-time)

Stack + delivery

PyTorch, Librosa DSP, 2D CNN, Bi-LSTM, and Attention.

PyTorchLibrosaPython
+15.8% accuracy over baseline

Build graph pipeline

The work connects across infrastructure, backend, AI, and product execution.

Hover over nodes to explore upstream sources and downstream effects. Observe data flows moving through the architecture.

Infrastructure

Airflow

Orchestrates batch ETL processing & workflows.

Spark

Tunes data partitioning & processing of 5TB+/day.

Kafka

Ingests real-time telemetry & system event streams.

Aurora

High-performance relational DB storage.

Kubernetes

Schedules pods, autoscaling, and service meshes.

AWS S3

Stores raw parquet & json streams for batch processing.

AWS EKS

Managed Kubernetes for scalable spark & model services.

Backend

FastAPI

Delivers async endpoints with sub-100ms P95 latency.

Spring Boot

Powering enterprise microservices & Java logic.

gRPC

High-speed internal microservice RPC communication.

WebSocket

Handles real-time bi-directional server sync.

NGINX

Load balancing, reverse proxies & entry routing.

REST APIs

Designed endpoints with rate-limiting & auto-validation.

Applied AI

LangGraph

Multi-agent autonomous systems using state graphs.

RAG systems

Vector search retrieval pipelines to reduce hallucination.

PyTorch

Deep learning models for forecasting & classification.

Vector Search

High-dimensionality index matching on MongoDB/Atlas.

Model Delivery

Packaging & deploying weights to online APIs.

NLP / TF-IDF

Feature extraction, plot synopsis tokenization.

Scikit-Learn

Random Forest popularity classifiers & regressions.

Product

Hackathons

Rapid prototyping under high-pressure windows.

Agentic UX

Interfaces built for generative AI and agent states.

Storytelling

Making deep technical concepts readable & compelling.

Execution Speed

Optimizing startup, page transitions & DB calls.

Real-Time Systems

Building responsive tools with immediate feedback.

Streamlit HUD

Interactive dashboards for data quality & execution.

Power BI UI

Business insights and popularity analytics reports.

Pipeline Integration Context
$ hover a node to inspect system integrations...

GitHub pulse

Code is a daily habit, not a seasonal event.

Contribution heatmap from the last 6 months. Consistency is the strongest signal of a working builder.

Open GitHub

Contributions

847

PRs Merged

142

Total Stars

84

Global Rank

Top 4.2%

Current Streak

12d

Longest Streak

34d

Active Days

32

Core Libs

LangGraph

Mon
Wed
Fri
Less
More

Local Telemetry Sandbox

Kernel Packet Router

Click on the circle nodes (Center, Left, or Right) to toggle their routing direction!

THROUGHPUT: 0
HIGH SCORE: 0

diagnostics engine

Load Balancer Ingest

A simple packets pipeline test simulation. Change router states dynamically to match database, cache, and service requests.

Database request: route Left
Cache request: route Center
API request: route Right
kernel_space.oshero0%