16GB GPU ਉੱਤੇ ਸਥਾਨਕ AI ਕੋਡਿੰਗ: Strata, OpenCode ਅਤੇ Qwen3.8

Table of Contents
16GB GPU ਉੱਤੇ ਸਥਾਨਕ ਕੋਡਿੰਗ ਏਜੰਟ ਸੀਮਿਤ development ਕੰਮਾਂ ਲਈ ਇੱਕ ਵਿਹਾਰਕ ਵਿਕਲਪ ਹੈ। Strata ਅਤੇ OpenCode ਸਥਾਨਕ inference ਨੂੰ file editing, shell commands ਅਤੇ test execution ਨਾਲ ਜੋੜਦੇ ਹਨ। RTX 4060 Ti ਉੱਤੇ ਦੱਸੇ Qwen3.8-Flash-Next run ਨੇ ਭੁਗਤਾਨੀ inference calls ਤੋਂ ਬਿਨਾਂ application, ਬਾਅਦ ਦੀਆਂ ਤਬਦੀਲੀਆਂ ਅਤੇ racing-game prototype ਪੂਰਾ ਕੀਤਾ।
Subscription ਬਦਲਣ ਲਈ ਵੱਡਾ test ਲੋੜੀਂਦਾ ਹੈ। ਪ੍ਰਕਾਸ਼ਿਤ ਨਤੀਜੇ ਇੱਕ machine ਉੱਤੇ ਚੱਲਦੀ setup ਦਿਖਾਉਂਦੇ ਹਨ। ਇਹ ਵੱਖ-ਵੱਖ languages, repositories ਜਾਂ ਔਖੇ debugging ਕੰਮਾਂ ਵਿੱਚ paid coding services ਨਾਲ ਬਰਾਬਰੀ ਸਾਬਤ ਨਹੀਂ ਕਰਦੇ। Hardware ਵਿੱਚ 64GB system RAM ਵੀ ਹੈ, ਜਿਸ ਵਿੱਚ model ਦਾ ਵੱਡਾ ਹਿੱਸਾ ਰਹਿੰਦਾ ਹੈ।
ਮੁੱਖ ਗੱਲਾਂ
- 16GB GPU memory ਦੱਸਦਾ ਹੈ, ਦੱਸੀ configuration ਲਈ ਲੋੜੀਂਦੀ ਕੁੱਲ memory ਨਹੀਂ।
- Prompt reuse ਮਹੱਤਵਪੂਰਨ ਹੈ, ਕਿਉਂਕਿ coding agents ਵਾਰ-ਵਾਰ overlap ਕਰਦੀ conversation history ਭੇਜਦੇ ਹਨ।
- Reasoning ਲਈ budget ਚਾਹੀਦਾ ਹੈ, ਤਾਂ planning code ਅਤੇ tool calls ਲਈ ਥਾਂ ਛੱਡੇ।
- Generated tests pass ਹੋਣਾ ਅਧੂਰਾ ਸਬੂਤ ਹੈ, Docker ਅਤੇ gameplay ਲਈ ਵੱਖਰੀ ਜਾਂਚ ਚਾਹੀਦੀ ਹੈ।
- Zero API spend ownership costs ਨੂੰ ਨਹੀਂ ਗਿਣਦਾ, ਜਿਵੇਂ ਬਿਜਲੀ ਅਤੇ maintenance time।
Prerequisites: Compatible computer, ਚੁਣੇ model ਲਈ ਕਾਫ਼ੀ free storage, Git, npm ਸਮੇਤ Node.js ਅਤੇ terminal tools ਦੀ ਜਾਣਕਾਰੀ। ਦੱਸੀ IQ3_S setup ਲਈ 64GB RAM ਵਰਤੋ। ਹੋਰ configurations ਲਈ current Strata installer ਵੇਖੋ।
ਸਮਾਂ ਅਤੇ ਮੁਸ਼ਕਲ: ਦਰਮਿਆਨਾ। Task completion speed ਮਾਪਣ ਤੋਂ ਪਹਿਲਾਂ ਵੱਡੇ model downloads ਅਤੇ installation ਲਈ ਸਮਾਂ ਰੱਖੋ। ਦੱਸੇ task timings setup ਨੂੰ ਸ਼ਾਮਲ ਨਹੀਂ ਕਰਦੇ।
Test ਕੀਤੀ configuration
| Component | ਦੱਸੀ configuration |
|---|---|
| GPU | NVIDIA RTX 4060 Ti, 16GB VRAM |
| CPU | Intel Core i5-11600K |
| System memory | 64GB RAM |
| Storage | NVMe SSD |
| Model | Qwen3.8-Flash-Next, 125B MoE, IQ3_S |
| Inference engine | Strata |
| Coding agent | OpenCode |
| Configured context | 65,536 tokens |
| Configured output limit | 16,384 tokens |
NetworkCoder ਦੀ published configuration ਅਤੇ results ਇਹ ਅੰਕੜੇ ਦਿੰਦੇ ਹਨ। Test repository 46.84GiB expert weights ਨੂੰ 28 seconds ਵਿੱਚ load ਕਰਨ ਦੀ report ਕਰਦੀ ਹੈ। 4,431 experts ਨੇ GPU memory ਦੀ 8.45GiB ਵਰਤੀ। ਇਹ measurements tested run ਨੂੰ ਦੱਸਦੇ ਹਨ, ਕਿਸੇ ਹੋਰ engine version ਉੱਤੇ guaranteed allocation ਨੂੰ ਨਹੀਂ।
Current compatibility ਇਸ test ਤੋਂ ਵੱਡੀ ਹੈ। 6 October 2026 ਨੂੰ ਕੀਤੀ ਜਾਂਚ ਮੁਤਾਬਕ Strata project ਘੱਟੋ-ਘੱਟ 12GB VRAM ਵਾਲੇ ਚੁਣੇ NVIDIA ਅਤੇ AMD cards, ਨਾਲੇ 32GB RAM systems ਲਈ ਛੋਟੇ models document ਕਰਦਾ ਹੈ। ਇਹ options 16GB GPU, 64GB RAM ਅਤੇ IQ3_S ਵਾਲੇ experiment ਨੂੰ ਦੁਹਰਾਉਂਦੇ ਨਹੀਂ। Hardware ਖਰੀਦਣ ਤੋਂ ਪਹਿਲਾਂ exact GPU, model variant ਅਤੇ runtime support check ਕਰੋ।
Model ਕਿੱਥੇ ਰਹਿੰਦਾ ਹੈ
Mixture of experts, ਜਾਂ MoE, ਹਰ token ਲਈ model ਦੀਆਂ expert networks ਦਾ ਇੱਕ subset activate ਕਰਦਾ ਹੈ। ਇਸ ਨਾਲ ਹਰ step ਉੱਤੇ ਹਰ parameter ਵਰਤਣ ਦੇ ਮੁਕਾਬਲੇ active computation ਘੱਟ ਹੁੰਦੀ ਹੈ। ਬਾਕੀ weights ਨੂੰ storage ਅਤੇ computation ਤੱਕ ਰਸਤਾ ਫਿਰ ਵੀ ਚਾਹੀਦਾ ਹੈ।
Strata ਦੀ hybrid execution ਅਕਸਰ ਵਰਤੇ ਜਾਣ ਵਾਲੇ experts ਨੂੰ GPU ਉੱਤੇ ਰੱਖਦੀ ਹੈ ਅਤੇ expert collection ਨੂੰ system RAM ਵਿੱਚ ਰੱਖਦੀ ਹੈ। CPU uncached experts ਨੂੰ ਉੱਥੇ compute ਕਰਦਾ ਹੈ ਅਤੇ GPU cached experts ਨੂੰ ਸੰਭਾਲਦਾ ਹੈ। Storage model files ਅਤੇ lookup data ਲਈ ਵੀ ਵਰਤੀ ਜਾਂਦੀ ਹੈ। System ਪੂਰੇ 125B model ਨੂੰ 16GB VRAM ਵਿੱਚ ਨਹੀਂ ਸਮਾਉਂਦਾ।
GPU memory ਦੇ ਕਈ ਮੁਕਾਬਲਾਤੀ ਕੰਮ ਹਨ। Runtime allocations, attention state ਅਤੇ expert cache ਸੀਮਿਤ resource ਸਾਂਝਾ ਕਰਦੇ ਹਨ। Context ਲਈ ਵੱਧ ਜਗ੍ਹਾ expert-cache capacity ਘਟਾਉਂਦੀ ਹੈ। Current Strata versions KV-cache streaming ਵੀ support ਕਰਦੀਆਂ ਹਨ, ਇਸ ਲਈ exact placement ਪੁਰਾਣੀ configuration ਤੋਂ ਵੱਖਰੀ ਹੈ। ਆਪਣੇ engine version ਲਈ technical documentation ਵੇਖੋ।

GPU inference system ਦਾ ਇੱਕ ਹਿੱਸਾ ਹੈ, CPU execution, system RAM ਅਤੇ storage ਦੇ ਨਾਲ
Prompt reuse speed ਬਦਲਦਾ ਹੈ
Coding agent ਇੱਕ loop ਚਲਾਉਂਦਾ ਹੈ: task ਪੜ੍ਹਦਾ ਹੈ, tool action ਮੰਗਦਾ ਹੈ, result ਲੈਂਦਾ ਹੈ ਅਤੇ ਅਗਲਾ action ਚੁਣਦਾ ਹੈ। File contents, errors ਅਤੇ test output conversation ਵਿੱਚ ਇਕੱਠੇ ਹੁੰਦੇ ਹਨ। Agents context ਨੂੰ compact ਜਾਂ select ਵੀ ਕਰਦੇ ਹਨ, ਇਸ ਲਈ ਹਰ implementation ਬਿਨਾਂ ਬਦਲੀ ਪੂਰੀ history ਹਮੇਸ਼ਾ ਦੁਬਾਰਾ ਨਹੀਂ ਭੇਜਦੀ।
Prefill generation ਤੋਂ ਪਹਿਲਾਂ input tokens process ਕਰਦਾ ਹੈ। Decode response ਬਣਾਉਂਦਾ ਹੈ। Fast decode ਪਰ slow repeated prefill ਵਾਲਾ system ਵੀ tool calls ਦੇ ਵਿਚਕਾਰ ਉਡੀਕ ਕਰਵਾਉਂਦਾ ਹੈ।
ਦੱਸੇ 44K-token timing ਵਿੱਚ prefix reuse ਵਰਤੀ ਗਈ। Strata ਨੇ ਪਹਿਲਾਂ process ਕੀਤੀ conversation content ਮੁੜ ਵਰਤੀ ਅਤੇ ਵਧਦਾ prompt ਲਗਭਗ ਇੱਕ ਤੋਂ ਤਿੰਨ seconds ਵਿੱਚ ਤਿਆਰ ਸੀ। ਇਹ continuing agent session ਲਈ ਲਾਭਦਾਇਕ evidence ਹੈ। ਇਹ 44,000 ਪੂਰੀ ਤਰ੍ਹਾਂ ਨਵੇਂ tokens ਨੂੰ scratch ਤੋਂ ਇੱਕ second ਵਿੱਚ process ਕਰਨ ਦਾ ਸਬੂਤ ਨਹੀਂ।
| Measurement | ਕੀ record ਕਰਨਾ ਹੈ |
|---|---|
| Cold prompt | Reusable prefix state ਤੋਂ ਬਿਨਾਂ ਨਵੀਂ content ਦਾ ਸਮਾਂ |
| Warm continuation | Existing context ਵਿੱਚ tool result ਜੋੜਨ ਤੋਂ ਬਾਅਦ ਦਾ ਸਮਾਂ |
| Generation speed | Response ਦੌਰਾਨ tokens per second |
| Task duration | Planning, generation, tools, tests ਅਤੇ retries ਇਕੱਠੇ |
ਦੋ caches ਵੱਖਰੇ ਕੰਮ ਕਰਦੇ ਹਨ। Expert cache ਅਕਸਰ ਵਰਤੇ weights ਨੂੰ GPU computation ਦੇ ਨੇੜੇ ਰੱਖਦਾ ਹੈ। Prefix reuse ਪੁਰਾਣੇ input ਉੱਤੇ ਕੰਮ ਦੁਹਰਾਉਣ ਤੋਂ ਬਚਾਉਂਦਾ ਹੈ। High expert-cache hit rate prompt-cache hit ਸਾਬਤ ਨਹੀਂ ਕਰਦਾ।
Tasks ਨੇ ਕੀ ਦਿਖਾਇਆ
| Task | ਦੱਸਿਆ ਸਮਾਂ | ਦੱਸਿਆ ਨਤੀਜਾ |
|---|---|---|
| Build task manager | 3m 38s | Express API, interface, 5/5 tests |
| Fix editing and add dates | 4m 58s | Changes complete, 6/6 tests |
| Add export/import and packaging | 3m 48s | 7/7 tests, Docker build unverified |
| Racing game, first attempt | 6m 35s | Reasoning ਦੌਰਾਨ output ਖਤਮ, code ਨਹੀਂ |
| Racing game, budgeted retry | 4m 51s | Game generated, JavaScript syntax checked |
Published results file ਪਹਿਲੇ task ਲਈ 48–52 tokens per second, ਦੂਜੇ ਲਈ 40–43 ਅਤੇ ਤੀਜੇ ਲਈ 37–44 output speed ਦਰਜ ਕਰਦੀ ਹੈ। “Up to 52” ਇਹਨਾਂ observations ਵਿੱਚ peak ਹੈ, ਹਰ task ਲਈ sustained rate ਨਹੀਂ।
ਪਹਿਲੇ ਤਿੰਨ tasks ਇੱਕ application ਨੂੰ ਵਧਾਉਂਦੇ ਹਨ। 5/5, 6/6 ਅਤੇ 7/7 counts successive test suites ਦੱਸਦੇ ਹਨ। ਇਹਨਾਂ ਨੂੰ ਜੋੜਨਾ 18 independent capabilities ਸਾਬਤ ਨਹੀਂ ਕਰਦਾ। ਉਸੇ agent ਵੱਲੋਂ ਲਿਖੇ tests ਨੂੰ coverage ਅਤੇ meaningful assertions ਲਈ review ਚਾਹੀਦਾ ਹੈ।
Environmental recovery ਕੰਮ ਦਾ ਹਿੱਸਾ ਸੀ। Agent ਨੇ inappropriate shell command ਤੋਂ recover ਕੀਤਾ ਅਤੇ testing ਦੌਰਾਨ stale application server ਪਛਾਣਿਆ। ਇਹ ਲਾਭਦਾਇਕ behaviours ਹਨ, ਪਰ existing process terminate ਕਰਨ ਲਈ shared development environment ਵਿੱਚ ਸੋਚੇ-ਸਮਝੇ permissions ਚਾਹੀਦੇ ਹਨ।
Docker unverified ਰਿਹਾ। Agent ਨੇ Dockerfile ਲਿਖਿਆ, ਪਰ build ਲਈ running Docker engine ਨਹੀਂ ਸੀ। Container ਤੋਂ ਬਾਹਰ application tests pass ਹੋਣ ਨਾਲ working image ਸਾਬਤ ਨਹੀਂ ਹੁੰਦੀ। Game retry ਨੇ syntax check ਕੀਤੀ ਅਤੇ browser ਖੋਲ੍ਹਿਆ, ਪਰ agent ਕੋਲ gameplay ਦੀ direct visual confirmation ਨਹੀਂ ਸੀ।
Output ਲਈ ਥਾਂ ਰੱਖੋ
{
"reasoning_budget_tokens": 8000
}
ਇਹ setting ਮੌਜੂਦਾ strata-iq3_s.json configuration ਵਿੱਚ merge ਕਰੋ, ਬਾਕੀ fields ਬਚਾ ਕੇ, ਫਿਰ ਚੁਣੇ Strata model ਨੂੰ restart ਕਰੋ। ਇਹ Strata setting ਹੈ, replacement OpenCode configuration file ਨਹੀਂ। Startup output ਵਿੱਚ active budget verify ਕਰੋ।
Strata ਇੱਕ hard reasoning budget document ਕਰਦਾ ਹੈ, ਜੋ thinking phase ਖਤਮ ਕਰਕੇ answer ਵੱਲ ਜਾਂਦਾ ਹੈ। Request-level value configured default ਨੂੰ override ਕਰਦੀ ਹੈ। ਇਹ setting low ਜਾਂ high ਵਰਗੀ general reasoning-effort instruction ਤੋਂ ਵੱਖਰੀ ਹੈ।
ਪਹਿਲੀ game attempt ਨੇ planning ਦੌਰਾਨ ਆਪਣਾ 16,384-token allowance ਖਰਚ ਕਰ ਦਿੱਤਾ। Retry ਨੇ 8,000-token thinking budget ਲਗਾਇਆ ਅਤੇ code ਦਿੱਤਾ। ਇਹ workload ਲਈ bounded reasoning phase ਵਰਤਣ ਦੀ ਹਮਾਇਤ ਕਰਦਾ ਹੈ। ਇਹ ਹਰ task ਲਈ 8,000 ਨੂੰ best setting ਸਾਬਤ ਨਹੀਂ ਕਰਦਾ।
ਬਾਕੀ output allowance maximum ਹੈ, guarantee ਕੀਤੀ reservation ਨਹੀਂ। ਜੇ 16,384 total limit ਹੇਠ 8,000-token cap ਪੂਰਾ ਵਰਤਿਆ ਜਾਂਦਾ, ਤਾਂ ਹੋਰ overhead ਤੋਂ ਪਹਿਲਾਂ ਲਗਭਗ 8,384 tokens ਬਚਦੇ। Result log retry ਉੱਤੇ 11,054 written tokens ਦੱਸਦਾ ਹੈ, ਪਰ reasoning ਅਤੇ code ਦੀ complete ਵੰਡ ਨਹੀਂ ਦਿੰਦਾ। ਇਸ ਨੂੰ ਉਸੇ limit ਵਿੱਚ 8,000 reasoning tokens plus 11,054 code tokens ਨਾ ਸਮਝੋ।
Narrow edits ਲਈ ਛੋਟਾ budget ਵਰਤੋ, ਫਿਰ ਵਧੇਰੇ analysis ਵਾਲੇ tasks ਲਈ ਵੱਡੇ budgets test ਕਰੋ। Complete file output, finish status ਅਤੇ test results ਵੇਖੋ। ਵੱਧ planning time ਤਦ ਹੀ ਲਾਭਦਾਇਕ ਹੈ ਜਦੋਂ delivered change ਸੁਧਰੇ।
Strata ਅਤੇ OpenCode ਜੋੜੋ
git clone https://github.com/Niko1221/Strata.git
cd Strata
./setup.sh
ਇਹ Linux setup entry point ਹੈ। Current installation instructions review ਕਰੋ ਅਤੇ documented Windows path ਲਈ START-HERE.bat ਵਰਤੋ। Reported configuration ਦੇ ਨੇੜੇ ਜਾਣ ਲਈ original Qwen3.8-Flash-Next IQ3_S variant ਅਤੇ 65,536-token context ਚੁਣੋ। Benchmark notes ਵਿੱਚ engine version ਅਤੇ selected model files save ਕਰੋ।
npm install -g opencode-ai
OpenCode install ਕਰੋ
official instructions
ਨਾਲ। ਵੱਖਰੇ terminal ਵਿੱਚ STRATA_BASE_URL ਨੂੰ Strata ਵੱਲੋਂ ਦਿੱਤੀ local API base ਵਜੋਂ define ਕਰੋ, ਜਿਸ ਵਿੱਚ /v1 suffix ਹੋਵੇ। ਇਸ setup ਲਈ inference ਨੂੰ local machine ਨਾਲ ਬੰਨ੍ਹ ਕੇ ਰੱਖੋ।
{
"provider": {
"strata": {
"npm": "@ai-sdk/openai-compatible",
"name": "Strata local",
"options": {
"baseURL": "{env:STRATA_BASE_URL}",
"apiKey": "local"
},
"models": {
"qwen3.8-flash-next-iq3_s": {
"name": "Qwen3.8-Flash-Next IQ3_S",
"limit": { "context": 65536, "output": 16384 }
}
}
}
},
"model": "strata/qwen3.8-flash-next-iq3_s"
}
Provider block ਨੂੰ ~/.config/opencode/opencode.json ਵਿੱਚ merge ਕਰੋ, existing settings ਬਚਾ ਕੇ। ਇਹ
published example configuration
ਨੂੰ adapt ਕਰਦਾ ਹੈ ਅਤੇ local endpoint environment variable ਤੋਂ ਲੋਡ ਕਰਦਾ ਹੈ। OpenCode
custom providers
ਅਤੇ
environment substitution
document ਕਰਦਾ ਹੈ।
local placeholder credential ਹੈ, example unauthenticated local setup ਦੇ ਅਨੁਸਾਰ। ਇਹ server secure ਨਹੀਂ ਕਰਦਾ। ਜੇ Strata instance authentication enable ਕਰਦੀ ਹੈ, ਤਾਂ configured credential ਨੂੰ ਉਚਿਤ local secret mechanism ਰਾਹੀਂ ਦਿਓ।
opencode ਨੂੰ disposable project copy ਵਿੱਚ start ਕਰੋ ਅਤੇ configured model ਚੁਣੋ। Code ਭੇਜਣ ਤੋਂ ਪਹਿਲਾਂ selected provider verify ਕਰੋ। Client-side context declaration server ਦੇ configured context ਨੂੰ ਨਹੀਂ ਵਧਾਉਂਦੀ, ਅਤੇ 65,536-token window input ਲਈ 65,536 tokens ਨਹੀਂ ਛੱਡਦੀ ਜਦੋਂ output ਲਈ ਵੀ ਥਾਂ ਚਾਹੀਦੀ ਹੈ।
Local inference ਅਤੇ permissions
{
"permission": {
"edit": "ask",
"bash": "ask"
}
}
Agent evaluate ਕਰਦੇ ਸਮੇਂ ਇਹ initial permissions OpenCode configuration ਵਿੱਚ merge ਕਰੋ। Permission documentation available controls ਸਮਝਾਉਂਦੀ ਹੈ। File changes ਅਤੇ shell execution inference ਜਿੱਥੇ ਵੀ ਚੱਲੇ, ਤੁਹਾਡੀ machine ਨੂੰ ਪ੍ਰਭਾਵਿਤ ਕਰਦੇ ਹਨ।
Local inference ਹਰ tool ਨੂੰ local ਨਹੀਂ ਬਣਾਉਂਦੀ। Package downloads, web tools, external integrations ਅਤੇ optional sharing ਵਿੱਚ network services ਸ਼ਾਮਲ ਰਹਿੰਦੀਆਂ ਹਨ। Private repositories ਸੰਭਾਲਣ ਤੋਂ ਪਹਿਲਾਂ enabled tools review ਕਰੋ। “Model locally runs” ਦਾ ਦਾਅਵਾ “ਕੁਝ ਵੀ computer ਤੋਂ ਬਾਹਰ ਨਹੀਂ ਜਾਂਦਾ” ਨਾਲੋਂ ਛੋਟਾ ਹੈ।
ਪਹਿਲੀ run ਦੀ troubleshooting
| Symptom | ਪਹਿਲਾਂ ਕੀ check ਕਰਨਾ ਹੈ |
|---|---|
| Provider unavailable | Strata ਚੱਲ ਰਹੀ ਹੈ ਅਤੇ environment variable OpenCode process ਤੱਕ ਪਹੁੰਚਦੀ ਹੈ |
| Model selection ਵਿੱਚ missing | Provider ID ਅਤੇ model ID saved configuration ਨਾਲ match ਕਰਦੇ ਹਨ |
| Context-limit error | Prompt length plus requested output server ਦੀ active window ਵਿੱਚ fit ਹੁੰਦਾ ਹੈ |
| Planning without code | Reasoning budget, total output limit ਅਤੇ finish status |
| Slow continuation | Prefix reuse, memory pressure, expert-cache behaviour ਅਤੇ competing processes |
| Docker command fails | Working engine ਉਪਲਬਧ ਹੈ, ਕੇਵਲ command-line client ਨਹੀਂ |
ਇੱਕ ਵਾਰ ਵਿੱਚ ਇੱਕ setting ਬਦਲੋ ਅਤੇ saved starting state ਤੋਂ ਉਹੀ task ਦੁਹਰਾਓ। ਇਸ ਨਾਲ configuration improvement ਨੂੰ ਵੱਖਰੇ prompt ਜਾਂ ਆਸਾਨ test ਤੋਂ ਵੱਖ ਕੀਤਾ ਜਾ ਸਕਦਾ ਹੈ।
ਕੀ ਇਹ subscription ਬਦਲਦਾ ਹੈ?
| Local setup test ਕਰਨ ਦੀ ਸਥਿਤੀ | ਹੋਰ option ਉਪਲਬਧ ਰੱਖੋ |
|---|---|
| Existing compatible hardware | ਕੇਵਲ untested workload ਲਈ hardware ਖਰੀਦਣਾ |
| Bounded application changes | ਵੱਡੀਆਂ ਅਣਜਾਣ repositories ਅਤੇ ਔਖੀਆਂ migrations |
| Repeatable acceptance tests | Correctness check ਕਰਨ ਦੇ reliable ਤਰੀਕੇ ਤੋਂ ਬਿਨਾਂ tasks |
| Runtime maintenance ਲਈ ਸਮਾਂ | Minimal setup ਅਤੇ support overhead ਵਾਲਾ ਕੰਮ |
Zero API spend ਮਹੱਤਵਪੂਰਨ ਹੈ, ਖਾਸ ਕਰਕੇ ਜਦੋਂ hardware ਪਹਿਲਾਂ ਉਪਲਬਧ ਹੋਵੇ। ਇਸ ਵਿੱਚ electricity, hardware depreciation, storage ਅਤੇ environment maintain ਕਰਨ ਦਾ time ਸ਼ਾਮਲ ਨਹੀਂ। Displayed zero-cost field ਇਹ expenses ਵੀ meter ਨਹੀਂ ਕਰਦੀ।
Subscription comparison ਲਈ matched tasks ਚਾਹੀਦੇ ਹਨ। ਦੋਵੇਂ systems ਵਿੱਚ ਉਹੀ starting repository, instructions ਅਤੇ acceptance checks ਚਲਾਓ। Elapsed time ਨਾਲ retries ਅਤੇ human corrections ਵੀ record ਕਰੋ। Full workflow evaluate ਕਰਦੇ ਸਮੇਂ failed first game attempt ਸ਼ਾਮਲ ਕਰੋ, ਕੇਵਲ successful retry ਨਹੀਂ।
Useful local agent ਨੂੰ universal superiority ਦੀ ਲੋੜ ਨਹੀਂ। ਜੇ ਇਹ routine edits ਭਰੋਸੇ ਨਾਲ ਕਰਦਾ ਹੈ ਅਤੇ ਔਖੇ tasks ਦਾ ਛੋਟਾ set ਕਿਸੇ ਹੋਰ tool ਨੂੰ ਛੱਡਦਾ ਹੈ, ਤਾਂ ਇਹ ਪਹਿਲਾਂ ਹੀ ਤੁਹਾਡੀ paid services ਦੀ ਲੋੜ ਬਦਲਦਾ ਹੈ। ਆਪਣੇ projects ਵਿੱਚ accepted work ਦੇ ਆਧਾਰ ਉੱਤੇ ਫੈਸਲਾ ਕਰੋ।
Demonstration ਅਤੇ ਅਗਲੇ ਕਦਮ
ਹੋਰ ਵੇਖੋ: 125B local coding-agent demonstration । ਉੱਪਰ ਦਿੱਤੇ measurements published test ਨਾਲ ਸਬੰਧਤ ਹਨ, ਇਸ article ਲਈ ਕੀਤੇ independent benchmark ਨਾਲ ਨਹੀਂ।
- ਇੱਕ ਛੋਟਾ task ਦੁਹਰਾਓ fixed acceptance criteria ਅਤੇ saved starting revision ਨਾਲ।
- Cold ਅਤੇ warm turns ਮਾਪੋ, prefix reuse ਨੂੰ cold prefill speed ਨਾ ਸਮਝੋ।
- Packaging ਨੂੰ ਵੱਖਰੇ ਤੌਰ ਉੱਤੇ verify ਕਰੋ successful image build, startup ਅਤੇ container-level checks ਨਾਲ।
- Generated tests inspect ਕਰੋ ਅਤੇ ਉਹ cases ਜੋ implementation ਨੇ anticipate ਨਹੀਂ ਕੀਤੇ, ਜੋੜੋ।
- Completed work compare ਕਰੋ subscription ਬਦਲਣ ਤੋਂ ਪਹਿਲਾਂ ਆਪਣੇ current coding tool ਨਾਲ।
Hardware planning ਲਈ, local AI model ਅਤੇ GPU context guide ਪੜ੍ਹੋ। Hosted-model behaviour ਲਈ OpenRouter provider routing ਅਤੇ costs ਵੇਖੋ।






