Table of Contents

16GB GPU ਉੱਤੇ ਸਥਾਨਕ ਕੋਡਿੰਗ ਏਜੰਟ ਸੀਮਿਤ development ਕੰਮਾਂ ਲਈ ਇੱਕ ਵਿਹਾਰਕ ਵਿਕਲਪ ਹੈ। Strata ਅਤੇ OpenCode ਸਥਾਨਕ inference ਨੂੰ file editing, shell commands ਅਤੇ test execution ਨਾਲ ਜੋੜਦੇ ਹਨ। RTX 4060 Ti ਉੱਤੇ ਦੱਸੇ Qwen3.8-Flash-Next run ਨੇ ਭੁਗਤਾਨੀ inference calls ਤੋਂ ਬਿਨਾਂ application, ਬਾਅਦ ਦੀਆਂ ਤਬਦੀਲੀਆਂ ਅਤੇ racing-game prototype ਪੂਰਾ ਕੀਤਾ।

Subscription ਬਦਲਣ ਲਈ ਵੱਡਾ test ਲੋੜੀਂਦਾ ਹੈ। ਪ੍ਰਕਾਸ਼ਿਤ ਨਤੀਜੇ ਇੱਕ machine ਉੱਤੇ ਚੱਲਦੀ setup ਦਿਖਾਉਂਦੇ ਹਨ। ਇਹ ਵੱਖ-ਵੱਖ languages, repositories ਜਾਂ ਔਖੇ debugging ਕੰਮਾਂ ਵਿੱਚ paid coding services ਨਾਲ ਬਰਾਬਰੀ ਸਾਬਤ ਨਹੀਂ ਕਰਦੇ। Hardware ਵਿੱਚ 64GB system RAM ਵੀ ਹੈ, ਜਿਸ ਵਿੱਚ model ਦਾ ਵੱਡਾ ਹਿੱਸਾ ਰਹਿੰਦਾ ਹੈ।

ਮੁੱਖ ਗੱਲਾਂ

  • 16GB GPU memory ਦੱਸਦਾ ਹੈ, ਦੱਸੀ configuration ਲਈ ਲੋੜੀਂਦੀ ਕੁੱਲ memory ਨਹੀਂ।
  • Prompt reuse ਮਹੱਤਵਪੂਰਨ ਹੈ, ਕਿਉਂਕਿ coding agents ਵਾਰ-ਵਾਰ overlap ਕਰਦੀ conversation history ਭੇਜਦੇ ਹਨ।
  • Reasoning ਲਈ budget ਚਾਹੀਦਾ ਹੈ, ਤਾਂ planning code ਅਤੇ tool calls ਲਈ ਥਾਂ ਛੱਡੇ।
  • Generated tests pass ਹੋਣਾ ਅਧੂਰਾ ਸਬੂਤ ਹੈ, Docker ਅਤੇ gameplay ਲਈ ਵੱਖਰੀ ਜਾਂਚ ਚਾਹੀਦੀ ਹੈ।
  • Zero API spend ownership costs ਨੂੰ ਨਹੀਂ ਗਿਣਦਾ, ਜਿਵੇਂ ਬਿਜਲੀ ਅਤੇ maintenance time।

Prerequisites: Compatible computer, ਚੁਣੇ model ਲਈ ਕਾਫ਼ੀ free storage, Git, npm ਸਮੇਤ Node.js ਅਤੇ terminal tools ਦੀ ਜਾਣਕਾਰੀ। ਦੱਸੀ IQ3_S setup ਲਈ 64GB RAM ਵਰਤੋ। ਹੋਰ configurations ਲਈ current Strata installer ਵੇਖੋ।

ਸਮਾਂ ਅਤੇ ਮੁਸ਼ਕਲ: ਦਰਮਿਆਨਾ। Task completion speed ਮਾਪਣ ਤੋਂ ਪਹਿਲਾਂ ਵੱਡੇ model downloads ਅਤੇ installation ਲਈ ਸਮਾਂ ਰੱਖੋ। ਦੱਸੇ task timings setup ਨੂੰ ਸ਼ਾਮਲ ਨਹੀਂ ਕਰਦੇ।

Test ਕੀਤੀ configuration

Componentਦੱਸੀ configuration
GPUNVIDIA RTX 4060 Ti, 16GB VRAM
CPUIntel Core i5-11600K
System memory64GB RAM
StorageNVMe SSD
ModelQwen3.8-Flash-Next, 125B MoE, IQ3_S
Inference engineStrata
Coding agentOpenCode
Configured context65,536 tokens
Configured output limit16,384 tokens

NetworkCoder ਦੀ published configuration ਅਤੇ results ਇਹ ਅੰਕੜੇ ਦਿੰਦੇ ਹਨ। Test repository 46.84GiB expert weights ਨੂੰ 28 seconds ਵਿੱਚ load ਕਰਨ ਦੀ report ਕਰਦੀ ਹੈ। 4,431 experts ਨੇ GPU memory ਦੀ 8.45GiB ਵਰਤੀ। ਇਹ measurements tested run ਨੂੰ ਦੱਸਦੇ ਹਨ, ਕਿਸੇ ਹੋਰ engine version ਉੱਤੇ guaranteed allocation ਨੂੰ ਨਹੀਂ।

Current compatibility ਇਸ test ਤੋਂ ਵੱਡੀ ਹੈ। 6 October 2026 ਨੂੰ ਕੀਤੀ ਜਾਂਚ ਮੁਤਾਬਕ Strata project ਘੱਟੋ-ਘੱਟ 12GB VRAM ਵਾਲੇ ਚੁਣੇ NVIDIA ਅਤੇ AMD cards, ਨਾਲੇ 32GB RAM systems ਲਈ ਛੋਟੇ models document ਕਰਦਾ ਹੈ। ਇਹ options 16GB GPU, 64GB RAM ਅਤੇ IQ3_S ਵਾਲੇ experiment ਨੂੰ ਦੁਹਰਾਉਂਦੇ ਨਹੀਂ। Hardware ਖਰੀਦਣ ਤੋਂ ਪਹਿਲਾਂ exact GPU, model variant ਅਤੇ runtime support check ਕਰੋ।

Model ਕਿੱਥੇ ਰਹਿੰਦਾ ਹੈ

Mixture of experts, ਜਾਂ MoE, ਹਰ token ਲਈ model ਦੀਆਂ expert networks ਦਾ ਇੱਕ subset activate ਕਰਦਾ ਹੈ। ਇਸ ਨਾਲ ਹਰ step ਉੱਤੇ ਹਰ parameter ਵਰਤਣ ਦੇ ਮੁਕਾਬਲੇ active computation ਘੱਟ ਹੁੰਦੀ ਹੈ। ਬਾਕੀ weights ਨੂੰ storage ਅਤੇ computation ਤੱਕ ਰਸਤਾ ਫਿਰ ਵੀ ਚਾਹੀਦਾ ਹੈ।

Strata ਦੀ hybrid execution ਅਕਸਰ ਵਰਤੇ ਜਾਣ ਵਾਲੇ experts ਨੂੰ GPU ਉੱਤੇ ਰੱਖਦੀ ਹੈ ਅਤੇ expert collection ਨੂੰ system RAM ਵਿੱਚ ਰੱਖਦੀ ਹੈ। CPU uncached experts ਨੂੰ ਉੱਥੇ compute ਕਰਦਾ ਹੈ ਅਤੇ GPU cached experts ਨੂੰ ਸੰਭਾਲਦਾ ਹੈ। Storage model files ਅਤੇ lookup data ਲਈ ਵੀ ਵਰਤੀ ਜਾਂਦੀ ਹੈ। System ਪੂਰੇ 125B model ਨੂੰ 16GB VRAM ਵਿੱਚ ਨਹੀਂ ਸਮਾਉਂਦਾ।

GPU memory ਦੇ ਕਈ ਮੁਕਾਬਲਾਤੀ ਕੰਮ ਹਨ। Runtime allocations, attention state ਅਤੇ expert cache ਸੀਮਿਤ resource ਸਾਂਝਾ ਕਰਦੇ ਹਨ। Context ਲਈ ਵੱਧ ਜਗ੍ਹਾ expert-cache capacity ਘਟਾਉਂਦੀ ਹੈ। Current Strata versions KV-cache streaming ਵੀ support ਕਰਦੀਆਂ ਹਨ, ਇਸ ਲਈ exact placement ਪੁਰਾਣੀ configuration ਤੋਂ ਵੱਖਰੀ ਹੈ। ਆਪਣੇ engine version ਲਈ technical documentation ਵੇਖੋ।

ਸਥਾਨਕ workstation ਵੱਲੋਂ model execution ਨੂੰ graphics card, system memory ਅਤੇ solid-state storage ਵਿੱਚ ਵੰਡਣ ਦੀ ਤਸਵੀਰ

GPU inference system ਦਾ ਇੱਕ ਹਿੱਸਾ ਹੈ, CPU execution, system RAM ਅਤੇ storage ਦੇ ਨਾਲ

Prompt reuse speed ਬਦਲਦਾ ਹੈ

Coding agent ਇੱਕ loop ਚਲਾਉਂਦਾ ਹੈ: task ਪੜ੍ਹਦਾ ਹੈ, tool action ਮੰਗਦਾ ਹੈ, result ਲੈਂਦਾ ਹੈ ਅਤੇ ਅਗਲਾ action ਚੁਣਦਾ ਹੈ। File contents, errors ਅਤੇ test output conversation ਵਿੱਚ ਇਕੱਠੇ ਹੁੰਦੇ ਹਨ। Agents context ਨੂੰ compact ਜਾਂ select ਵੀ ਕਰਦੇ ਹਨ, ਇਸ ਲਈ ਹਰ implementation ਬਿਨਾਂ ਬਦਲੀ ਪੂਰੀ history ਹਮੇਸ਼ਾ ਦੁਬਾਰਾ ਨਹੀਂ ਭੇਜਦੀ।

Prefill generation ਤੋਂ ਪਹਿਲਾਂ input tokens process ਕਰਦਾ ਹੈ। Decode response ਬਣਾਉਂਦਾ ਹੈ। Fast decode ਪਰ slow repeated prefill ਵਾਲਾ system ਵੀ tool calls ਦੇ ਵਿਚਕਾਰ ਉਡੀਕ ਕਰਵਾਉਂਦਾ ਹੈ।

ਦੱਸੇ 44K-token timing ਵਿੱਚ prefix reuse ਵਰਤੀ ਗਈ। Strata ਨੇ ਪਹਿਲਾਂ process ਕੀਤੀ conversation content ਮੁੜ ਵਰਤੀ ਅਤੇ ਵਧਦਾ prompt ਲਗਭਗ ਇੱਕ ਤੋਂ ਤਿੰਨ seconds ਵਿੱਚ ਤਿਆਰ ਸੀ। ਇਹ continuing agent session ਲਈ ਲਾਭਦਾਇਕ evidence ਹੈ। ਇਹ 44,000 ਪੂਰੀ ਤਰ੍ਹਾਂ ਨਵੇਂ tokens ਨੂੰ scratch ਤੋਂ ਇੱਕ second ਵਿੱਚ process ਕਰਨ ਦਾ ਸਬੂਤ ਨਹੀਂ।

Measurementਕੀ record ਕਰਨਾ ਹੈ
Cold promptReusable prefix state ਤੋਂ ਬਿਨਾਂ ਨਵੀਂ content ਦਾ ਸਮਾਂ
Warm continuationExisting context ਵਿੱਚ tool result ਜੋੜਨ ਤੋਂ ਬਾਅਦ ਦਾ ਸਮਾਂ
Generation speedResponse ਦੌਰਾਨ tokens per second
Task durationPlanning, generation, tools, tests ਅਤੇ retries ਇਕੱਠੇ

ਦੋ caches ਵੱਖਰੇ ਕੰਮ ਕਰਦੇ ਹਨ। Expert cache ਅਕਸਰ ਵਰਤੇ weights ਨੂੰ GPU computation ਦੇ ਨੇੜੇ ਰੱਖਦਾ ਹੈ। Prefix reuse ਪੁਰਾਣੇ input ਉੱਤੇ ਕੰਮ ਦੁਹਰਾਉਣ ਤੋਂ ਬਚਾਉਂਦਾ ਹੈ। High expert-cache hit rate prompt-cache hit ਸਾਬਤ ਨਹੀਂ ਕਰਦਾ।

Tasks ਨੇ ਕੀ ਦਿਖਾਇਆ

Taskਦੱਸਿਆ ਸਮਾਂਦੱਸਿਆ ਨਤੀਜਾ
Build task manager3m 38sExpress API, interface, 5/5 tests
Fix editing and add dates4m 58sChanges complete, 6/6 tests
Add export/import and packaging3m 48s7/7 tests, Docker build unverified
Racing game, first attempt6m 35sReasoning ਦੌਰਾਨ output ਖਤਮ, code ਨਹੀਂ
Racing game, budgeted retry4m 51sGame generated, JavaScript syntax checked

Published results file ਪਹਿਲੇ task ਲਈ 48–52 tokens per second, ਦੂਜੇ ਲਈ 40–43 ਅਤੇ ਤੀਜੇ ਲਈ 37–44 output speed ਦਰਜ ਕਰਦੀ ਹੈ। “Up to 52” ਇਹਨਾਂ observations ਵਿੱਚ peak ਹੈ, ਹਰ task ਲਈ sustained rate ਨਹੀਂ।

ਪਹਿਲੇ ਤਿੰਨ tasks ਇੱਕ application ਨੂੰ ਵਧਾਉਂਦੇ ਹਨ। 5/5, 6/6 ਅਤੇ 7/7 counts successive test suites ਦੱਸਦੇ ਹਨ। ਇਹਨਾਂ ਨੂੰ ਜੋੜਨਾ 18 independent capabilities ਸਾਬਤ ਨਹੀਂ ਕਰਦਾ। ਉਸੇ agent ਵੱਲੋਂ ਲਿਖੇ tests ਨੂੰ coverage ਅਤੇ meaningful assertions ਲਈ review ਚਾਹੀਦਾ ਹੈ।

Environmental recovery ਕੰਮ ਦਾ ਹਿੱਸਾ ਸੀ। Agent ਨੇ inappropriate shell command ਤੋਂ recover ਕੀਤਾ ਅਤੇ testing ਦੌਰਾਨ stale application server ਪਛਾਣਿਆ। ਇਹ ਲਾਭਦਾਇਕ behaviours ਹਨ, ਪਰ existing process terminate ਕਰਨ ਲਈ shared development environment ਵਿੱਚ ਸੋਚੇ-ਸਮਝੇ permissions ਚਾਹੀਦੇ ਹਨ।

Docker unverified ਰਿਹਾ। Agent ਨੇ Dockerfile ਲਿਖਿਆ, ਪਰ build ਲਈ running Docker engine ਨਹੀਂ ਸੀ। Container ਤੋਂ ਬਾਹਰ application tests pass ਹੋਣ ਨਾਲ working image ਸਾਬਤ ਨਹੀਂ ਹੁੰਦੀ। Game retry ਨੇ syntax check ਕੀਤੀ ਅਤੇ browser ਖੋਲ੍ਹਿਆ, ਪਰ agent ਕੋਲ gameplay ਦੀ direct visual confirmation ਨਹੀਂ ਸੀ।

Output ਲਈ ਥਾਂ ਰੱਖੋ

{
  "reasoning_budget_tokens": 8000
}

ਇਹ setting ਮੌਜੂਦਾ strata-iq3_s.json configuration ਵਿੱਚ merge ਕਰੋ, ਬਾਕੀ fields ਬਚਾ ਕੇ, ਫਿਰ ਚੁਣੇ Strata model ਨੂੰ restart ਕਰੋ। ਇਹ Strata setting ਹੈ, replacement OpenCode configuration file ਨਹੀਂ। Startup output ਵਿੱਚ active budget verify ਕਰੋ।

Strata ਇੱਕ hard reasoning budget document ਕਰਦਾ ਹੈ, ਜੋ thinking phase ਖਤਮ ਕਰਕੇ answer ਵੱਲ ਜਾਂਦਾ ਹੈ। Request-level value configured default ਨੂੰ override ਕਰਦੀ ਹੈ। ਇਹ setting low ਜਾਂ high ਵਰਗੀ general reasoning-effort instruction ਤੋਂ ਵੱਖਰੀ ਹੈ।

ਪਹਿਲੀ game attempt ਨੇ planning ਦੌਰਾਨ ਆਪਣਾ 16,384-token allowance ਖਰਚ ਕਰ ਦਿੱਤਾ। Retry ਨੇ 8,000-token thinking budget ਲਗਾਇਆ ਅਤੇ code ਦਿੱਤਾ। ਇਹ workload ਲਈ bounded reasoning phase ਵਰਤਣ ਦੀ ਹਮਾਇਤ ਕਰਦਾ ਹੈ। ਇਹ ਹਰ task ਲਈ 8,000 ਨੂੰ best setting ਸਾਬਤ ਨਹੀਂ ਕਰਦਾ।

ਬਾਕੀ output allowance maximum ਹੈ, guarantee ਕੀਤੀ reservation ਨਹੀਂ। ਜੇ 16,384 total limit ਹੇਠ 8,000-token cap ਪੂਰਾ ਵਰਤਿਆ ਜਾਂਦਾ, ਤਾਂ ਹੋਰ overhead ਤੋਂ ਪਹਿਲਾਂ ਲਗਭਗ 8,384 tokens ਬਚਦੇ। Result log retry ਉੱਤੇ 11,054 written tokens ਦੱਸਦਾ ਹੈ, ਪਰ reasoning ਅਤੇ code ਦੀ complete ਵੰਡ ਨਹੀਂ ਦਿੰਦਾ। ਇਸ ਨੂੰ ਉਸੇ limit ਵਿੱਚ 8,000 reasoning tokens plus 11,054 code tokens ਨਾ ਸਮਝੋ।

Narrow edits ਲਈ ਛੋਟਾ budget ਵਰਤੋ, ਫਿਰ ਵਧੇਰੇ analysis ਵਾਲੇ tasks ਲਈ ਵੱਡੇ budgets test ਕਰੋ। Complete file output, finish status ਅਤੇ test results ਵੇਖੋ। ਵੱਧ planning time ਤਦ ਹੀ ਲਾਭਦਾਇਕ ਹੈ ਜਦੋਂ delivered change ਸੁਧਰੇ।

Strata ਅਤੇ OpenCode ਜੋੜੋ

git clone https://github.com/Niko1221/Strata.git
cd Strata
./setup.sh

ਇਹ Linux setup entry point ਹੈ। Current installation instructions review ਕਰੋ ਅਤੇ documented Windows path ਲਈ START-HERE.bat ਵਰਤੋ। Reported configuration ਦੇ ਨੇੜੇ ਜਾਣ ਲਈ original Qwen3.8-Flash-Next IQ3_S variant ਅਤੇ 65,536-token context ਚੁਣੋ। Benchmark notes ਵਿੱਚ engine version ਅਤੇ selected model files save ਕਰੋ।

npm install -g opencode-ai

OpenCode install ਕਰੋ official instructions ਨਾਲ। ਵੱਖਰੇ terminal ਵਿੱਚ STRATA_BASE_URL ਨੂੰ Strata ਵੱਲੋਂ ਦਿੱਤੀ local API base ਵਜੋਂ define ਕਰੋ, ਜਿਸ ਵਿੱਚ /v1 suffix ਹੋਵੇ। ਇਸ setup ਲਈ inference ਨੂੰ local machine ਨਾਲ ਬੰਨ੍ਹ ਕੇ ਰੱਖੋ।

{
  "provider": {
    "strata": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Strata local",
      "options": {
        "baseURL": "{env:STRATA_BASE_URL}",
        "apiKey": "local"
      },
      "models": {
        "qwen3.8-flash-next-iq3_s": {
          "name": "Qwen3.8-Flash-Next IQ3_S",
          "limit": { "context": 65536, "output": 16384 }
        }
      }
    }
  },
  "model": "strata/qwen3.8-flash-next-iq3_s"
}

Provider block ਨੂੰ ~/.config/opencode/opencode.json ਵਿੱਚ merge ਕਰੋ, existing settings ਬਚਾ ਕੇ। ਇਹ published example configuration ਨੂੰ adapt ਕਰਦਾ ਹੈ ਅਤੇ local endpoint environment variable ਤੋਂ ਲੋਡ ਕਰਦਾ ਹੈ। OpenCode custom providers ਅਤੇ environment substitution document ਕਰਦਾ ਹੈ।

local placeholder credential ਹੈ, example unauthenticated local setup ਦੇ ਅਨੁਸਾਰ। ਇਹ server secure ਨਹੀਂ ਕਰਦਾ। ਜੇ Strata instance authentication enable ਕਰਦੀ ਹੈ, ਤਾਂ configured credential ਨੂੰ ਉਚਿਤ local secret mechanism ਰਾਹੀਂ ਦਿਓ।

opencode ਨੂੰ disposable project copy ਵਿੱਚ start ਕਰੋ ਅਤੇ configured model ਚੁਣੋ। Code ਭੇਜਣ ਤੋਂ ਪਹਿਲਾਂ selected provider verify ਕਰੋ। Client-side context declaration server ਦੇ configured context ਨੂੰ ਨਹੀਂ ਵਧਾਉਂਦੀ, ਅਤੇ 65,536-token window input ਲਈ 65,536 tokens ਨਹੀਂ ਛੱਡਦੀ ਜਦੋਂ output ਲਈ ਵੀ ਥਾਂ ਚਾਹੀਦੀ ਹੈ।

Local inference ਅਤੇ permissions

{
  "permission": {
    "edit": "ask",
    "bash": "ask"
  }
}

Agent evaluate ਕਰਦੇ ਸਮੇਂ ਇਹ initial permissions OpenCode configuration ਵਿੱਚ merge ਕਰੋ। Permission documentation available controls ਸਮਝਾਉਂਦੀ ਹੈ। File changes ਅਤੇ shell execution inference ਜਿੱਥੇ ਵੀ ਚੱਲੇ, ਤੁਹਾਡੀ machine ਨੂੰ ਪ੍ਰਭਾਵਿਤ ਕਰਦੇ ਹਨ।

Local inference ਹਰ tool ਨੂੰ local ਨਹੀਂ ਬਣਾਉਂਦੀ। Package downloads, web tools, external integrations ਅਤੇ optional sharing ਵਿੱਚ network services ਸ਼ਾਮਲ ਰਹਿੰਦੀਆਂ ਹਨ। Private repositories ਸੰਭਾਲਣ ਤੋਂ ਪਹਿਲਾਂ enabled tools review ਕਰੋ। “Model locally runs” ਦਾ ਦਾਅਵਾ “ਕੁਝ ਵੀ computer ਤੋਂ ਬਾਹਰ ਨਹੀਂ ਜਾਂਦਾ” ਨਾਲੋਂ ਛੋਟਾ ਹੈ।

ਪਹਿਲੀ run ਦੀ troubleshooting

Symptomਪਹਿਲਾਂ ਕੀ check ਕਰਨਾ ਹੈ
Provider unavailableStrata ਚੱਲ ਰਹੀ ਹੈ ਅਤੇ environment variable OpenCode process ਤੱਕ ਪਹੁੰਚਦੀ ਹੈ
Model selection ਵਿੱਚ missingProvider ID ਅਤੇ model ID saved configuration ਨਾਲ match ਕਰਦੇ ਹਨ
Context-limit errorPrompt length plus requested output server ਦੀ active window ਵਿੱਚ fit ਹੁੰਦਾ ਹੈ
Planning without codeReasoning budget, total output limit ਅਤੇ finish status
Slow continuationPrefix reuse, memory pressure, expert-cache behaviour ਅਤੇ competing processes
Docker command failsWorking engine ਉਪਲਬਧ ਹੈ, ਕੇਵਲ command-line client ਨਹੀਂ

ਇੱਕ ਵਾਰ ਵਿੱਚ ਇੱਕ setting ਬਦਲੋ ਅਤੇ saved starting state ਤੋਂ ਉਹੀ task ਦੁਹਰਾਓ। ਇਸ ਨਾਲ configuration improvement ਨੂੰ ਵੱਖਰੇ prompt ਜਾਂ ਆਸਾਨ test ਤੋਂ ਵੱਖ ਕੀਤਾ ਜਾ ਸਕਦਾ ਹੈ।

ਕੀ ਇਹ subscription ਬਦਲਦਾ ਹੈ?

Local setup test ਕਰਨ ਦੀ ਸਥਿਤੀਹੋਰ option ਉਪਲਬਧ ਰੱਖੋ
Existing compatible hardwareਕੇਵਲ untested workload ਲਈ hardware ਖਰੀਦਣਾ
Bounded application changesਵੱਡੀਆਂ ਅਣਜਾਣ repositories ਅਤੇ ਔਖੀਆਂ migrations
Repeatable acceptance testsCorrectness check ਕਰਨ ਦੇ reliable ਤਰੀਕੇ ਤੋਂ ਬਿਨਾਂ tasks
Runtime maintenance ਲਈ ਸਮਾਂMinimal setup ਅਤੇ support overhead ਵਾਲਾ ਕੰਮ

Zero API spend ਮਹੱਤਵਪੂਰਨ ਹੈ, ਖਾਸ ਕਰਕੇ ਜਦੋਂ hardware ਪਹਿਲਾਂ ਉਪਲਬਧ ਹੋਵੇ। ਇਸ ਵਿੱਚ electricity, hardware depreciation, storage ਅਤੇ environment maintain ਕਰਨ ਦਾ time ਸ਼ਾਮਲ ਨਹੀਂ। Displayed zero-cost field ਇਹ expenses ਵੀ meter ਨਹੀਂ ਕਰਦੀ।

Subscription comparison ਲਈ matched tasks ਚਾਹੀਦੇ ਹਨ। ਦੋਵੇਂ systems ਵਿੱਚ ਉਹੀ starting repository, instructions ਅਤੇ acceptance checks ਚਲਾਓ। Elapsed time ਨਾਲ retries ਅਤੇ human corrections ਵੀ record ਕਰੋ। Full workflow evaluate ਕਰਦੇ ਸਮੇਂ failed first game attempt ਸ਼ਾਮਲ ਕਰੋ, ਕੇਵਲ successful retry ਨਹੀਂ।

Useful local agent ਨੂੰ universal superiority ਦੀ ਲੋੜ ਨਹੀਂ। ਜੇ ਇਹ routine edits ਭਰੋਸੇ ਨਾਲ ਕਰਦਾ ਹੈ ਅਤੇ ਔਖੇ tasks ਦਾ ਛੋਟਾ set ਕਿਸੇ ਹੋਰ tool ਨੂੰ ਛੱਡਦਾ ਹੈ, ਤਾਂ ਇਹ ਪਹਿਲਾਂ ਹੀ ਤੁਹਾਡੀ paid services ਦੀ ਲੋੜ ਬਦਲਦਾ ਹੈ। ਆਪਣੇ projects ਵਿੱਚ accepted work ਦੇ ਆਧਾਰ ਉੱਤੇ ਫੈਸਲਾ ਕਰੋ।

Demonstration ਅਤੇ ਅਗਲੇ ਕਦਮ

ਹੋਰ ਵੇਖੋ: 125B local coding-agent demonstration । ਉੱਪਰ ਦਿੱਤੇ measurements published test ਨਾਲ ਸਬੰਧਤ ਹਨ, ਇਸ article ਲਈ ਕੀਤੇ independent benchmark ਨਾਲ ਨਹੀਂ।

  1. ਇੱਕ ਛੋਟਾ task ਦੁਹਰਾਓ fixed acceptance criteria ਅਤੇ saved starting revision ਨਾਲ।
  2. Cold ਅਤੇ warm turns ਮਾਪੋ, prefix reuse ਨੂੰ cold prefill speed ਨਾ ਸਮਝੋ।
  3. Packaging ਨੂੰ ਵੱਖਰੇ ਤੌਰ ਉੱਤੇ verify ਕਰੋ successful image build, startup ਅਤੇ container-level checks ਨਾਲ।
  4. Generated tests inspect ਕਰੋ ਅਤੇ ਉਹ cases ਜੋ implementation ਨੇ anticipate ਨਹੀਂ ਕੀਤੇ, ਜੋੜੋ।
  5. Completed work compare ਕਰੋ subscription ਬਦਲਣ ਤੋਂ ਪਹਿਲਾਂ ਆਪਣੇ current coding tool ਨਾਲ।

Hardware planning ਲਈ, local AI model ਅਤੇ GPU context guide ਪੜ੍ਹੋ। Hosted-model behaviour ਲਈ OpenRouter provider routing ਅਤੇ costs ਵੇਖੋ।