惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Schneier on Security
博客园_首页
量子位
博客园 - 司徒正美
S
SegmentFault 最新的问题
J
Java Code Geeks
小众软件
小众软件
博客园 - 【当耐特】
The Register - Security
The Register - Security
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Microsoft Azure Blog
Microsoft Azure Blog
G
Google Developers Blog
Blog — PlanetScale
Blog — PlanetScale
T
Tailwind CSS Blog
博客园 - Franky
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
G
GRAHAM CLULEY
Cyberwarzone
Cyberwarzone
腾讯CDC
Apple Machine Learning Research
Apple Machine Learning Research
V
Visual Studio Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
The Hacker News
The Hacker News
aimingoo的专栏
aimingoo的专栏
V
Vulnerabilities – Threatpost
P
Palo Alto Networks Blog
Scott Helme
Scott Helme
L
LINUX DO - 热门话题
F
Full Disclosure
D
DataBreaches.Net
Martin Fowler
Martin Fowler
Cisco Talos Blog
Cisco Talos Blog
L
LINUX DO - 最新话题
云风的 BLOG
云风的 BLOG
C
Check Point Blog
T
Threatpost
Google DeepMind News
Google DeepMind News
WordPress大学
WordPress大学
W
WeLiveSecurity
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
有赞技术团队
有赞技术团队
Hugging Face - Blog
Hugging Face - Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
L
Lohrmann on Cybersecurity
Last Week in AI
Last Week in AI
T
Tor Project blog
T
Troy Hunt's Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
S
Security Affairs
SecWiki News
SecWiki News

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Why I Move AI Model Calls to the Server — Security, Performance, and Everything In Between
David Essien · 2026-05-26 · via DEV Community

When I was building Logicvisor — an AI-powered tool that reviews your algorithmic code, breaks down time and space complexity, and gives you the kind of feedback you'd want before a technical interview — I had to make a foundational architectural decision early on.

Where do the AI calls actually live?

It sounds simple. It really isn't. And I think it's a decision a lot of developers make too quickly, usually defaulting to whatever gets something running fastest. So I want to walk through how I thought about it, what the tradeoffs actually look like in practice, and why for Logicvisor — and honestly most production projects I work on — the answer was never really up for debate.


Table of Contents

  • The Case for Client-Side AI Calls
  • Why Client-Side Breaks Down in Production
  • The Case for Server-Side AI Calls
  • How This Played Out in Logicvisor
    • Performance
    • Security
    • Authentication
    • Caching
    • Provider Abstraction
    • AI Response Processing
  • So Was It Ever Really a Debate?

First, Let's Establish What We're Even Choosing Between

When your app needs to talk to an AI model — Gemini, Claude, GPT, whatever — that HTTP request to the model provider has to originate somewhere. You have two options:

Client-side: The browser makes the call directly to the AI provider's API.

Server-side: The browser calls your server, your server calls the AI provider, and the response comes back through your infrastructure.

That's the whole decision. But the consequences of each branch run deep.


The Case for Client-Side AI Calls

Let's be fair to the other side first, because client-side AI calls aren't just laziness — there are legitimate reasons to reach for them.

Zero backend overhead. If you're prototyping, building an MVP, or hacking something together for a weekend project, standing up a server just to proxy AI calls adds friction you might not need yet. The client calls the API, gets a response, done.

One less network hop. Client → AI provider is a straight line. Client → your server → AI provider is also a straight line, but a longer one. Every additional hop is a potential source of latency, and if your server is not geographically close to the AI provider, that gap compounds.

Fast iteration during development. Tweak a prompt, refresh the page, see the result. No redeployment cycle, no server restart. For the early exploratory phase of building with AI, this feedback loop is genuinely valuable.

Fine for purely client-facing tools. If you're building something that doesn't touch your own database, doesn't need user sessions, and doesn't have sensitive business logic — a personal productivity tool, a browser extension, an internal utility — client-side calls can be perfectly appropriate.

So that's the honest upside. Now here's where it falls apart.


Why Client-Side AI Calls Break Down in Production

The API Key Problem

This is the most obvious one, but it's worth being precise about why it's as bad as it is.

When you make an API call from the browser, your API key has to be in that request. There's no way around this — the provider needs to authenticate you. And since that request is made from the browser, the key is accessible to anyone who opens DevTools, intercepts traffic, or extracts it from your bundled JavaScript.

The consequence isn't just that someone can see your key. It's that they can use it. At your expense. Without your knowledge. AI API billing is usage-based, which means a single bad actor with your key can run up a bill that drains your account before your monitoring even fires an alert — if you have monitoring at all.

Key rotation helps, but it's reactive. The damage is usually already done.

Your Prompt Engineering Is Public

This one gets less attention but matters more than people realize.

The prompts you write are often where your actual product value lives. If you've spent time crafting a system prompt that makes your AI reviewer give structured, consistent, high-quality feedback on algorithmic code — that prompt is the product. Client-side calls expose it completely. A competitor can open DevTools, read your system prompt, and replicate your core feature in an afternoon.

On the server, your prompts never leave your infrastructure. The client sends input; the server decides what to do with it.

You Have No Control Over Abuse

On the client side, there's nothing stopping a user from writing a script that hammers your AI endpoint in a loop. Every one of those requests hits the AI provider and costs you tokens. You have no rate limiting, no request validation, no way to enforce quotas per user.

You're not just vulnerable to malicious actors either — a bug in your own frontend code that causes unintended re-fetching can silently burn through your API budget.

No Caching

AI API calls cost money per token. If multiple users ask your tool to review functionally identical code, why would you want to pay for that same inference three hundred times?

On the client side, you can't cache at the API level. Every identical request goes to the provider, incurs latency, and costs tokens. On the server, you can cache responses intelligently — hash the input, check your cache layer, return the cached result. You pay once.

You're Flying Blind in Production

Without server-side infrastructure, you have no centralized view of how your AI layer is actually being used. Which prompts are performing well? Which inputs are producing garbage responses? Which users are hitting rate limits? Where is your token spend going?

Client-side AI calls mean you're guessing at all of this. Logs, monitoring, and observability — the basic instrumentation of a production system — require a server in the loop.


The Case for Server-Side AI Calls

With that context established, here's what you actually get when the AI calls live on the server.

The API Key Never Touches the Client

Your key lives in an environment variable on the server. The client has zero knowledge of it, zero access to it, and zero ability to extract it. This is the minimum acceptable security posture for any application that will see real users.

You Control Rate Limiting

You decide how many requests a given user can make in a given window. You can enforce this per account, per IP, per session — whatever your threat model calls for. Abuse becomes something you manage rather than something that happens to you.

export async function enforceAIRateLimit(
    userId: string,
    request?: NextRequest
): Promise<RateLimitResult> {
    const tierLimits = await getUserTierLimits(userId);

    if (!tierLimits) {
        throw new RateLimitError("Unable to determine user tier limits");
    }

    // Free/Pro users: per-minute limits
    if (tierLimits.ai_requests_per_minute !== null) {
        const minuteLimit = await checkRateLimit(userId, "ai_request", "minute");

        if (!minuteLimit.isAllowed) {
          await logRateLimitViolation(userId, "ai_request", "minute", ...);
          throw new RateLimitError(``Rate limit exceeded. Resets at ${minuteLimit.resetTime.toISOString()}``);
        }

        return minuteLimit;
    }

    // Admin users: daily + monthly
    const dailyLimit = await checkRateLimit(userId, "ai_request", "daily");
    // ...monthly check follows same pattern
}

Response Caching Becomes Possible

Identical or near-identical inputs can return cached results, cutting both latency and cost. For a tool like Logicvisor where multiple users might submit similar sorting algorithm implementations, the savings on repeated inferences compound quickly.

// Normalize the code to a canonical form before hashing
// so that formatting differences don't result in cache misses
const canonicalCode = await canonicalizeCodeAST(solution, preferred_language);
const canonicalHash = await createCanonicalHash(
  typeof canonicalCode === "string" ? canonicalCode : ""
);

// Check cache before hitting the AI provider
const cachedReview = await getCachedAIReview(
  preferred_model.id + "-" + canonicalHash
);
if (cachedReview) {
  return NextResponse.json(
    { success: true, data: cachedReview },
    { status: 201 }
  );
}

You Can Inject Server-Side Context Into Prompts

This is where things get architecturally interesting. The client sends raw input — code, a question, a request. But your server knows things the client doesn't: who the user is, what their history looks like, what tier they're on, what language they've selected, what results they've already received. All of that context can be injected into the prompt before it ever leaves your infrastructure.

The client can't fake or manipulate that context because it never touches it.

Full Observability

Every request is logged. Every response is traceable. You can monitor token usage, flag anomalous behavior, track which prompts produce the best results, and debug production issues with actual data. This is what running software in production looks like.

await updateAPIUsageAnalytics("/api/internal", true, responseTime, aiTokensUsed, estimatedCost, user.id, {
  modelName: modelUsed,
  modelProvider,
  inputTokens,
  outputTokens,
  totalTokens: aiTokensUsed,
  actualCostUsd,
});

The BFF Argument — One Round Trip Instead of Many

There's a broader architectural win here that goes beyond just AI calls. Without a server in the middle, the client has to orchestrate everything itself: call the AI, wait for the response, then maybe hit your database, wait again, then update the UI. Each of those is a visible pause for the user.

With a Backend for Frontend (BFF) pattern, the client makes one request. The server handles the AI call, processes the response, queries the database if needed, applies any business logic, and returns a single resolved payload. The user feels one network round trip instead of a cascading waterfall of them.


Architecture Diagrams

Before getting into the implementation specifics, here's the architectural difference visualised.

Client-Side AI Calls — The Problem

Client-side architecture diagram

Browser → AI Provider directly · API key exposed in transit

Server-Side AI Calls — The Solution

Server-side architecture diagram

Browser → API Route → [Cache · Rate Limiter · AI Provider · DB] → Browser


How This Played Out in Logicvisor

Let me get concrete. Here's what moving the AI layer to the server actually looked like in practice.

Performance — Collapsing the Waterfall

Logicvisor uses Supabase on the backend. At various points in the app, I need to pull data from multiple tables, run the AI review, and return everything the page needs in one shot.

If this were all happening on the client, you'd be looking at: call Supabase for user context → wait → call the AI provider → wait → call Supabase again for historical reviews → wait → render. Each of those waits is visible to the user, and each one is an opportunity for something to fail mid-chain.

On the server, those calls happen in close proximity to each other and to the data. The AI call, the database queries, and any necessary transformations all resolve server-side, and the client gets one clean response. The user experiences a single loading state, not a series of UI flickers.

There's also the matter of compute. Parsing and processing a large AI response — stripping JSON fences, validating structure, transforming the output into the format the UI expects — is work that browsers are not well-suited for. Browsers are memory and CPU constrained by design, and they're competing with the DOM, with other tabs, with everything the user has open. A server doesn't have those constraints.

Security — API Obfuscation as a Feature

Moving calls to the server means the client's entire interaction is scoped to your own API. It calls your endpoint, gets a response, done. It has no visibility into what your server does with that request internally — which external APIs it calls, what keys those calls carry, or how the response was constructed.

This isn't security through obscurity; the obfuscation is a structural property of the architecture. OS-level network tools could theoretically expose some of this, but you've dramatically raised the bar for what an attacker needs to do to compromise your stack.

Your system prompts — the part of the product that actually encodes your domain knowledge and review methodology — never leave the server. That's not a small thing.

Authentication — Stateless and Seamless

A server in the loop makes proper auth architecture dramatically cleaner. I was able to set HttpOnly cookies, attach signed JWT tokens, and build a stateless authentication and authorization system that the client participates in without controlling.

Without a server, you end up storing tokens in localStorage or client-side state, which is a well-documented attack surface. The session becomes something the client manages, which means it's something an attacker can manipulate.

Caching — Paying for Inference Once

Token costs are real. For Logicvisor, where users might submit variations of common algorithm patterns — bubble sort, binary search, dynamic programming problems — I can cache AI responses keyed on a normalized hash of the input. A user submitting a well-known algorithm implementation gets a fast, cached response. The AI provider gets called once.

This also improves response times for cached queries significantly. The round trip to an AI provider is the most expensive part of the request by a wide margin. Eliminating it for repeat queries is the single highest-leverage performance optimization available to you.

Provider Abstraction — Switching Models Without Touching the Client

Here's something the client-side approach makes nearly impossible: swapping AI providers without your frontend caring at all.

Logicvisor supports both Gemini and Groq depending on the user's selected model. Gemini for deeper analysis with its thinking budget and Google Search grounding, Groq for speed. Two different SDKs, two different response shapes, two different token counting strategies, two different pricing models. The client knows none of this. It sends a request, it gets a review back.

That abstraction only works because the AI calls live on the server:

switch (modelProvider.toLowerCase()) {
  case "google":
    // Gemini — thinking budget + Google Search grounding
    const response = await ai.models.generateContent({
      model: preferred_model.id,
      config: {
        thinkingConfig: { thinkingBudget: prompt.estimatedTokens },
        tools: [{ googleSearch: {} }],
        seed: SEED,
        temperature: TEMPERATURE,
      },
      contents,
    });
    aiReviewText = response.text ?? "";
    // Extract actual token counts from usageMetadata
    inputTokens = response.usageMetadata?.promptTokenCount || 0;
    outputTokens = response.usageMetadata?.candidatesTokenCount || 0;
    break;

  case "groq":
    // Llama 3.3 70B — forces structured JSON response
    const groqResponse = await groq.chat.completions.create({
      messages: [{ role: "user", content: prompt.content }],
      model: groqModel,
      temperature: TEMPERATURE,
      seed: SEED,
      response_format: { type: "json_object" },
      stream: false,
    });
    aiReviewText = groqResponse.choices[0].message.content ?? "";
    inputTokens = groqResponse.usage?.prompt_tokens || 0;
    outputTokens = groqResponse.usage?.completion_tokens || 0;
    break;
}

If I wanted to add a third provider tomorrow — say, Claude for a specific model tier — that's a new case block on the server. The client contract doesn't change. No frontend deployment, no API key exposure, no breaking changes for users mid-session.

Try doing this cleanly when the calls are in the browser. You'd be shipping provider-specific SDK logic, API keys, and token counting math directly to the client — and every provider swap would mean a frontend change. The server is the only place this kind of abstraction is clean.

This also matters for cost tracking. Notice that each branch extracts token counts differently — Gemini from usageMetadata, Groq from usage. That per-provider normalization feeds into the analytics pipeline downstream, giving you a consistent view of cost across providers regardless of how each SDK reports it. That's only possible because all the provider-specific handling is in one place.

AI Response Processing — Cleaning Up the Model's Output

Here's the unglamorous part that doesn't get written about enough: AI models don't always return clean output.

Gemini, for example, sometimes wraps JSON responses in markdown fences even when you've explicitly told it not to. If your UI is trying to parse that response and render structured data, you need to clean it before it gets anywhere near the client.

On the server, I handle all of that: strip the fences, validate the JSON structure, handle the cases where the model returned something unexpected, and only send a clean, predictable payload to the client. If something goes wrong at this layer, I can log it, inspect it, and fix the prompt. The client just sees a well-structured response or a proper error.

export function extractJSONFromMarkdown(markdownString: string) {
  try {
    // Remove the ```json wrapper
    let jsonString = markdownString.trim();

    // Remove ```json from start
    if (jsonString.startsWith("```json")) {
      jsonString = jsonString.substring(7);
    }

    // Remove ``` from end
    if (jsonString.endsWith("```")) {
      jsonString = jsonString.substring(0, jsonString.length - 3);
    }

    // Parse the JSON
    const parsedData: PromptResult = JSON.parse(jsonString.trim());

    // Extract and decode the markdown_review field
    let markdownReview = parsedData.markdown_review;
    if (markdownReview) {
      // Decode JSON escape sequences
      markdownReview = markdownReview
        .replace(/\\n/g, "\n")   // Convert \n to actual newlines
        .replace(/\\t/g, "\t")   // Convert \t to actual tabs
        .replace(/\\r/g, "\r")   // Convert \r to carriage returns
        .replace(/\\"/g, '"')    // Convert \" to actual quotes
        .replace(/\\\\/g, "\\"); // Convert \\ to actual backslashes

      return { parsedData, markdownReview };
    } else {
      console.log("No markdown_review field found");
      return null;
    }
  } catch (error) {
    console.error("Error parsing JSON:", error);
    return null;
  }
}

If this processing happened on the client, every user's browser would be doing it — inconsistently, with no visibility, and with no way to fix edge cases without a frontend deployment.


So Was It Ever Really a Debate?

Honestly, not for Logicvisor.

Client-side AI calls are a valid tool in specific, narrow contexts — personal tools, internal utilities, quick prototypes where security is not a concern and scale is not a goal. The moment you have real users, real API costs, and real data flowing through your system, the calculus changes completely.

Security, performance, observability, and maintainability all point in the same direction. The extra infrastructure is real overhead. But it's the kind of overhead that pays for itself the first time an abuse attempt hits your rate limiter instead of your API bill.

For any production system where AI inference is part of the core product — keep it on the server. Build the client thin. Let the backend do the heavy lifting.