Table of Contents

ਸਭ ਤੋਂ ਵਧੀਆ ਲੋਕਲ coding model ਉਹ ਹੈ ਜੋ ਤੁਹਾਡੇ project ਦਾ ਕਾਫ਼ੀ ਹਿੱਸਾ ਮੈਮੋਰੀ ਵਿੱਚ ਰੱਖੇ ਅਤੇ ਲਾਭਦਾਇਕ ਰਹਿਣ ਲਈ ਉਸਨੂੰ ਤੇਜ਼ੀ ਨਾਲ ਪੜ੍ਹੇ। Model size ਅਜੇ ਵੀ ਮਹੱਤਵ ਰੱਖਦਾ ਹੈ। ਜਿਹੜਾ ਮਾਡਲ system RAM ਵਿੱਚ data ਭੇਜਣ ਤੋਂ ਬਾਅਦ ਹੀ fit ਹੁੰਦਾ ਹੈ, ਉਹ stable context window ਵਾਲੇ ਛੋਟੇ ਮਾਡਲ ਨਾਲੋਂ ਅਕਸਰ ਮਾੜਾ ਮਹਿਸੂਸ ਹੁੰਦਾ ਹੈ।

Model ਅਤੇ memory budget ਇਕੱਠੇ ਚੁਣੋ। Tier list ਤੋਂ ਮਾਡਲ ਚੁਣ ਕੇ ਉਸਨੂੰ ਅਣਉਚਿਤ hardware ਉੱਤੇ ਨਾ ਥੋਪੋ।

ਇਹ guide coding agents ਉੱਤੇ ਕੇਂਦ੍ਰਿਤ ਹੈ, ਛੋਟੇ autocomplete prompts ਉੱਤੇ ਨਹੀਂ। Agent fix ਲਿਖਣ ਤੋਂ ਪਹਿਲਾਂ files, tool output, compiler errors ਅਤੇ test results ਪੜ੍ਹਦਾ ਹੈ। ਇਹ inputs context ਵਰਤਦੇ ਹਨ ਅਤੇ hardware ਦਾ ਫੈਸਲਾ ਬਦਲਦੇ ਹਨ।

ਵੀਡੀਓ ਇੱਕ ਹਵਾਲਾ ਹੈ

ਹੇਠਾਂ ਦਿੱਤੀ video ਕਈ GPU memory tiers ਵਿੱਚ local models ਦੀ ਤੁਲਨਾ ਕਰਦੀ ਹੈ। ਇਹ practical model-selection observations ਅਤੇ report ਕੀਤੇ hardware results ਲਈ ਲਾਭਦਾਇਕ ਹੈ। ਇਹ article memory behavior, agent context, prompt processing ਅਤੇ total ownership cost ਦੇ ਆਧਾਰ ਉੱਤੇ ਫੈਸਲਾ ਵੱਖਰੇ ਢੰਗ ਨਾਲ ਸੰਗਠਿਤ ਕਰਦਾ ਹੈ।

Tier lists ਕਿਉਂ ਫੇਲ੍ਹ ਹੁੰਦੀਆਂ ਹਨ

Model tier list ਆਮ ਤੌਰ ਉੱਤੇ parameter count ਨੂੰ GPU memory number ਨਾਲ ਜੋੜਦੀ ਹੈ। ਪਹਿਲੇ estimate ਲਈ ਇਹ shortcut ਲਾਭਦਾਇਕ ਹੈ। ਜਦੋਂ coding agent ਅਸਲੀ repository ਪੜ੍ਹਨਾ ਸ਼ੁਰੂ ਕਰਦਾ ਹੈ, ਇਹ ਲਾਭਦਾਇਕ ਨਹੀਂ ਰਹਿੰਦਾ।

Model weights ਪਹਿਲੀ allocation ਹਨ। KV cache ਵਧਣ ਵਾਲੀ allocation ਹੈ। ਇਹ active context ਦੀਆਂ attention keys ਅਤੇ values ਰੱਖਦਾ ਹੈ ਤਾਂ ਜੋ runtime ਹਰ generated token ਤੋਂ ਬਾਅਦ ਪੂਰੇ prompt ਨੂੰ ਮੁੜ calculate ਨਾ ਕਰੇ।

Memory consumerਇਸ ਵਿੱਚ ਕੀ ਰਹਿੰਦਾ ਹੈਇਹ ਕਿਉਂ ਮਹੱਤਵਪੂਰਨ ਹੈ
Model weightsQuantized parametersਇੱਕ ਵਾਰ ਦੀ loading requirement
KV cacheActive context ਦੀਆਂ notesPrompt ਅਤੇ conversation ਨਾਲ ਵਧਦਾ ਹੈ
Runtime buffersTemporary computation spaceBackend ਅਤੇ batch size ਅਨੁਸਾਰ ਬਦਲਦਾ ਹੈ
Agent instructionsSystem prompts ਅਤੇ tool definitionsProject files ਆਉਣ ਤੋਂ ਪਹਿਲਾਂ context ਵਰਤਦਾ ਹੈ

Model file ਦਾ card ਵਿੱਚ fit ਹੋਣਾ usable agent ਦਾ ਸਬੂਤ ਨਹੀਂ ਹੈ। Runtime ਨੂੰ cache, temporary buffers, tool definitions ਅਤੇ ਅਗਲੇ response ਲਈ ਥਾਂ ਚਾਹੀਦੀ ਹੈ।

Context ਅਸਲ budget ਹੈ

Coding agents context ਨੂੰ ਸਿਰਫ਼ source files ਉੱਤੇ ਨਹੀਂ ਖਰਚਦੇ। Budget ਵਿੱਚ system instructions, tool schemas, directory listings, shell output, compiler messages, test results ਅਤੇ ਪਿਛਲੀਆਂ conversation turns ਵੀ ਸ਼ਾਮਲ ਹਨ।

ਤਿੰਨ MCP connections ਦੀਆਂ verbose tool definitions agent ਵੱਲੋਂ project file ਖੋਲ੍ਹਣ ਤੋਂ ਪਹਿਲਾਂ ਕਈ ਹਜ਼ਾਰ tokens ਵਰਤ ਸਕਦੀਆਂ ਹਨ। ਛੋਟਾ default context ਫਿਰ code ਲਈ ਥੋੜ੍ਹੀ ਥਾਂ ਛੱਡਦਾ ਹੈ। Agent ਜਵਾਬ ਦਿੰਦਾ ਰਹਿੰਦਾ ਹੈ, ਪਰ multi-step repairs ਲਈ ਲੋੜੀਂਦਾ working set ਗੁਆ ਲੈਂਦਾ ਹੈ।

Ollama ਦੀ runtime FAQ 4,096 tokens ਦਾ default context ਦੱਸਦੀ ਹੈ। ਇਸਦੀ OLLAMA_CONTEXT_LENGTH setting default ਬਦਲਦੀ ਹੈ, ਜਦਕਿ OLLAMA_NUM_PARALLEL concurrent requests ਦੀ ਗਿਣਤੀ ਨਾਲ ਲੋੜੀਂਦੀ memory ਵਧਾਉਂਦੀ ਹੈ। Local benchmark ਨੂੰ single-request result ਨਾਲ compare ਕਰਨ ਤੋਂ ਪਹਿਲਾਂ ਦੋਵੇਂ values ਲਿਖੋ।

OLLAMA_CONTEXT_LENGTH=32768 OLLAMA_NUM_PARALLEL=1 ollama serve

ਇਹ example ਇੱਕ active request ਲਈ 32K default set ਕਰਦਾ ਹੈ। ਇਹ project files ਲਈ 32K tokens reserve ਨਹੀਂ ਕਰਦਾ। Instructions, tool definitions, input ਅਤੇ output ਹਾਲੇ ਵੀ window ਸਾਂਝੀ ਕਰਦੇ ਹਨ।

ਇੱਕ ਸਧਾਰਨ Context Record

GPUs ਦੀ ਤੁਲਨਾ ਤੋਂ ਪਹਿਲਾਂ ਇਹ values ਲਿਖੋ:

  • Project size: ਆਮ task ਵਿੱਚ files ਅਤੇ ਲਗਭਗ source lines।
  • Tool overhead: system prompt, MCP schemas, shell tools ਅਤੇ editor instructions।
  • Error payload: ਆਮ compiler ਅਤੇ test output ਦੀ ਲੰਬਾਈ।
  • Target context: truncation ਤੋਂ ਬਿਨਾਂ ਰੱਖਣਾ ਚਾਹੁੰਦੇ ਸਭ ਤੋਂ ਵੱਡੇ prompt ਦਾ ਆਕਾਰ।
  • Response allowance: planned patch ਅਤੇ explanation ਲਈ ਰਾਖਵੀਂ ਥਾਂ।

ਨਤੀਜੇ ਨੂੰ workload specification ਵਜੋਂ ਵਰਤੋ। ਇੱਕ file ਵਾਰ edit ਕਰਨ ਵਾਲੇ developer ਦੀ memory requirement ਉਸ developer ਤੋਂ ਵੱਖਰੀ ਹੈ ਜੋ agent ਨੂੰ monorepo ਵਿੱਚ API trace ਕਰਨ ਲਈ ਕਹਿੰਦਾ ਹੈ।

KV Cache ranking ਬਦਲਦਾ ਹੈ

ਸਮਾਨ parameter counts ਵਾਲੇ ਦੋ models ਦੀ context cost ਅਕਸਰ ਵੱਖਰੀ ਹੁੰਦੀ ਹੈ। ਹਰ layer ਨੂੰ ਵਧਦੇ cache ਵਿੱਚ ਯੋਗਦਾਨ ਦੇਣ ਵਾਲੇ dense model ਨੂੰ hybrid attention design ਵਾਲੇ model ਨਾਲੋਂ ਵੱਧ memory ਚਾਹੀਦੀ ਹੋ ਸਕਦੀ ਹੈ।

Reference material ਵਿੱਚ ਚਰਚਿਤ ਇੱਕ 27B coding model full attention layers ਦੀ ਸੀਮਿਤ ਗਿਣਤੀ ਵਰਤਦਾ ਹੈ, ਜਦਕਿ ਹੋਰ layers fixed-size summary ਵਰਤਦੀਆਂ ਹਨ। Report ਕੀਤੀ 128K context cost cache ਲਈ 9GB ਤੋਂ ਘੱਟ ਰਹਿੰਦੀ ਹੈ। ਹਰ layer ਵਿੱਚ full growing attention ਵਾਲੇ ਇਸੇ ਆਕਾਰ ਦੇ model ਨੂੰ ਉਸੇ context length ਉੱਤੇ report ਮੁਤਾਬਕ 21GB ਤੋਂ ਵੱਧ ਚਾਹੀਦੀ ਹੈ।

ਇਹ figures ਖਾਸ model architectures ਅਤੇ runtime settings ਦੱਸਦੇ ਹਨ। ਇਨ੍ਹਾਂ ਨੂੰ architecture inspect ਕਰਨ ਦਾ ਕਾਰਨ ਸਮਝੋ, universal memory formula ਨਹੀਂ।

Model behaviorContext effectHardware implication
Full attention at every layerਵੱਡਾ cache growthਲੰਮੇ prompts ਲਈ ਵੱਧ memory ਚਾਹੀਦੀ ਹੈ
Hybrid or sliding attentionਕੁਝ layers ਵਿੱਚ ਘੱਟ cache growthਉਸੇ model size ਉੱਤੇ ਵੱਧ context headroom
Mixture of expertsਹਰ token ਲਈ ਘੱਟ active parametersਹਰ token ਲਈ ਘੱਟ compute, ਪਰ total weights ਲਈ storage ਫਿਰ ਵੀ ਚਾਹੀਦੀ ਹੈ
Long-context extensionਵੱਡੀ working windowਵੱਧ cache memory ਅਤੇ prompt-processing work

Memory ਖਰੀਦਣ ਤੋਂ ਪਹਿਲਾਂ model architecture ਪੜ੍ਹੋ। Parameter count ਲੰਮੇ coding session ਦੀ cost ਲੁਕਾਉਂਦਾ ਹੈ।

GPU Memory Tier ਅਨੁਸਾਰ ਚੁਣੋ

ਚਾਰ ਤੋਂ ਅੱਠ GB

ਛੋਟੇ dense models practical choice ਰਹਿੰਦੇ ਹਨ। ਇਹ card ਵਿੱਚ fit ਹੁੰਦੇ ਹਨ, ਤੇਜ਼ ਜਵਾਬ ਦਿੰਦੇ ਹਨ ਅਤੇ autocomplete, short explanations ਅਤੇ narrow file edits ਲਈ ਚੰਗੇ ਹਨ।

CPU offload ਵਾਲਾ ਵੱਡਾ model text ਬਣਾ ਸਕਦਾ ਹੈ, ਪਰ response time ਅਕਸਰ ਸੀਮਾ ਬਣ ਜਾਂਦਾ ਹੈ। Coding agent ਨੂੰ ਵਾਰ-ਵਾਰ file reads ਅਤੇ tool calls ਕਰਨੇ ਪੈਂਦੇ ਹਨ। Four-token-per-second setup ਵਿੱਚ model technical ਤੌਰ ਉੱਤੇ ਚੱਲਣ ਦੇ ਬਾਵਜੂਦ ਹਰ repair ਲਈ ਲੰਮਾ wait ਹੁੰਦਾ ਹੈ।

Mixture-of-experts models ਹੋਰ ਰਸਤਾ ਦਿੰਦੇ ਹਨ। ਛੋਟਾ active ਹਿੱਸਾ compute pressure ਘਟਾਉਂਦਾ ਹੈ, ਜਦਕਿ ਪੂਰਾ weight set system memory ਵਿੱਚ ਅੰਸ਼ਕ ਤੌਰ ਉੱਤੇ ਰਹਿੰਦਾ ਹੈ। ਇਹ ਤਰੀਕਾ ample RAM ਅਤੇ fast transfer paths ਵਾਲੇ system ਨੂੰ ਲਾਭ ਦਿੰਦਾ ਹੈ।

WorkloadSuggested direction
AutocompleteShort prompt ਵਾਲਾ ਛੋਟਾ dense model
Single-file editsTool support ਵਾਲਾ ਛੋਟਾ instruct model
Repository-wide agentਪਹਿਲਾਂ ਵੱਡਾ GPU rent ਕਰੋ ਜਾਂ system memory ਵਧਾਓ
Private code under an NDALocal model ਵਰਤੋ, ਪਰ ਛੋਟਾ scope ਜਾਂ ਹੌਲੀ runs ਸਵੀਕਾਰੋ

ਇਸ tier ਵਿੱਚ ਵੱਡਾ model ਲੱਭਣ ਤੋਂ ਪਹਿਲਾਂ system RAM ਖਰੀਦੋ। Stable small model ਉਸ ਵੱਡੇ model ਨਾਲੋਂ ਚੰਗਾ ਹੈ ਜੋ ਜ਼ਿਆਦਾਤਰ ਸਮਾਂ bus ਉੱਤੇ data ਹਿਲਾਉਂਦਾ ਰਹਿੰਦਾ ਹੈ।

ਬਾਰ੍ਹਾਂ ਤੋਂ ਸੋਲ੍ਹਾਂ GB

ਇਹ tier 27B coding model ਲਈ ਰਸਤਾ ਖੋਲ੍ਹਦਾ ਹੈ, ਪਰ quantization choice ਕੇਂਦਰੀ ਬਣ ਜਾਂਦੀ ਹੈ। ਲਗਭਗ 17GB ਵਾਲਾ standard 4-bit build runtime overhead ਆਉਣ ਤੋਂ ਬਾਅਦ 16GB card ਵਿੱਚ fit ਨਹੀਂ ਹੁੰਦਾ।

ਲਗਭਗ 13GB ਵਾਲਾ 3-bit build context ਲਈ ਵੱਧ ਥਾਂ ਛੱਡਦਾ ਹੈ। 12GB card choice ਨੂੰ 2-bit build ਜਾਂ ਛੋਟੇ mixture-of-experts model ਵੱਲ ਧੱਕਦਾ ਹੈ। Quality ਖਾਸ quantizer ਅਤੇ calibration process ਉੱਤੇ ਨਿਰਭਰ ਕਰਦੀ ਹੈ, ਇਸ ਲਈ ਇੱਕੋ bit depth ਵਾਲੇ ਦੋ uploads ਅਕਸਰ ਵੱਖਰੇ coding results ਦਿੰਦੇ ਹਨ।

Quantized model download ਕਰਨ ਤੋਂ ਪਹਿਲਾਂ ਇਹ ਜਾਂਚੋ:

  • Quantizer author and release notes
  • Calibration data and evaluation results
  • Tokenizer compatibility
  • Tool-calling tests
  • Context length at the chosen quantization
  • Runtime support in Ollama, llama.cpp, or the selected front end

“2-bit” ਜਾਂ “3-bit” ਨੂੰ ਪੂਰੀ quality description ਨਾ ਸਮਝੋ। Packing method ਅਤੇ calibration record ਮਹੱਤਵ ਰੱਖਦੇ ਹਨ।

ਚੌਵੀ ਤੋਂ ਬੱਤੀ GB

ਇਹ 27B coding model ਲਈ ਸਭ ਤੋਂ flexible range ਹੈ। 24GB card ਅਕਸਰ useful context ਵਾਲਾ 4-bit build fit ਕਰਦੀ ਹੈ, ਪਰ full 128K window ਬਾਕੀ memory ਤੋਂ ਵੱਧ ਹੋ ਸਕਦੀ ਹੈ। 32GB card runtime ਨੂੰ cache ਅਤੇ temporary allocations ਲਈ ਵੱਧ ਥਾਂ ਦਿੰਦੀ ਹੈ।

ਇਹ range ownership ਨੂੰ justify ਕਰਨਾ ਵੀ ਸੌਖਾ ਬਣਾਉਂਦੀ ਹੈ। ਇੱਕ GPU two-card split ਦੀ complexity ਤੋਂ ਬਿਨਾਂ workload ਸੰਭਾਲਦਾ ਹੈ। System multi-GPU build ਨਾਲੋਂ ਘੱਟ power ਵਰਤਦਾ ਹੈ ਅਤੇ software support test ਕਰਨਾ ਸੌਖਾ ਹੁੰਦਾ ਹੈ।

CapacityPractical position
24GBContext limits ਮਾਪਣ ਵਾਲਾ strong 4-bit model
32GBBroader context headroom ਵਾਲਾ 4-bit ਜਾਂ 6-bit 27B model
48GBCore model ਬਦਲੇ ਬਿਨਾਂ higher precision ਜਾਂ larger cache

24GB ਤੋਂ 32GB range frequent private coding work ਲਈ sensible ownership zone ਹੈ। ਇਹ weakest offload behavior ਤੋਂ ਬਚਾਉਂਦੀ ਹੈ ਅਤੇ desktop case ਵਿੱਚ data-center hardware ਲਿਆਉਣ ਲਈ ਮਜਬੂਰ ਨਹੀਂ ਕਰਦੀ।

ਅਠਤਾਲੀ GB ਅਤੇ ਵੱਧ

ਵੱਧ memory ਦਾ ਅਰਥ ਆਪਣੇ ਆਪ ਨਵਾਂ model ਨਹੀਂ ਹੁੰਦਾ। ਉਹੀ 27B model 48GB card ਉੱਤੇ 8-bit precision, ਵੱਡੇ cache ਅਤੇ ਘੱਟ compromises ਨਾਲ ਚੱਲ ਸਕਦਾ ਹੈ। ਸੁਧਾਰ consistency, context room ਅਤੇ output quality ਵਿੱਚ ਹੁੰਦਾ ਹੈ, reasoning ਦੇ ਨਵੇਂ ਪੱਧਰ ਵਿੱਚ ਨਹੀਂ।

128GB ਉੱਤੇ ਫੈਸਲਾ ਬਦਲ ਜਾਂਦਾ ਹੈ। Much larger mixture-of-experts model ਸੰਭਵ ਹੁੰਦਾ ਹੈ, ਪਰ prompt processing ਗੰਭੀਰ concern ਬਣ ਜਾਂਦੀ ਹੈ। Model ਅਕਸਰ ਤੇਜ਼ੀ ਨਾਲ generate ਕਰਦਾ ਹੈ, ਜਦਕਿ ਵੱਡੀ repository ਜਾਂ fresh tool result ingest ਕਰਨ ਵਿੱਚ ਲੰਮਾ ਸਮਾਂ ਲੈਂਦਾ ਹੈ।

Large models ਲਈ prefill speed ਨੂੰ decode speed ਤੋਂ ਵੱਖਰਾ ਮਾਪੋ। Slow prompt read ਤੋਂ ਬਾਅਦ fast answer agent workflow ਵਿੱਚ ਫਿਰ ਵੀ slow ਮਹਿਸੂਸ ਹੁੰਦਾ ਹੈ।

Decode speed test ਦਾ ਸਿਰਫ਼ ਅੱਧਾ ਹਿੱਸਾ ਹੈ

Decode speed generated tokens per second ਮਾਪਦੀ ਹੈ। ਇਹ ਪੁੱਛਦੀ ਹੈ, “Model ਕਿੰਨੀ ਤੇਜ਼ੀ ਨਾਲ ਲਿਖਦਾ ਹੈ?” Prefill speed prompt processing ਮਾਪਦੀ ਹੈ। ਇਹ ਪੁੱਛਦੀ ਹੈ, “Model ਕਿੰਨੀ ਤੇਜ਼ੀ ਨਾਲ ਪੜ੍ਹਦਾ ਹੈ?”

Agent ਆਪਣਾ ਬਹੁਤ ਸਮਾਂ reading ਵਿੱਚ ਲਗਾਉਂਦਾ ਹੈ। ਹਰ tool call ਨਵਾਂ text ਜੋੜਦੀ ਹੈ। ਅਗਲਾ response ਸ਼ੁਰੂ ਹੋਣ ਤੋਂ ਪਹਿਲਾਂ ਲੰਮੀ source file, stack trace ਜਾਂ test log prompt ਵਿੱਚ ਦਾਖਲ ਹੁੰਦੀ ਹੈ।

MetricUser experience
Decode tokens per secondProcessing ਤੋਂ ਬਾਅਦ answer ਕਿੰਨੀ ਤੇਜ਼ੀ ਨਾਲ ਆਉਂਦਾ ਹੈ
Prompt tokens per secondAnswer ਸ਼ੁਰੂ ਹੋਣ ਤੋਂ ਪਹਿਲਾਂ agent ਕਿੰਨਾ wait ਕਰਦਾ ਹੈ
Time to first tokenPrompt processing ਅਤੇ setup ਦੀ ਮਿਲੀ ਹੋਈ delay
Context retentionTask ਦੌਰਾਨ project state ਦਾ ਕਿੰਨਾ ਹਿੱਸਾ ਉਪਲਬਧ ਰਹਿੰਦਾ ਹੈ

ਆਪਣੇ prompt sizes ਨੂੰ benchmark ਕਰੋ। Short synthetic prompt repository work ਦੌਰਾਨ ਸਭ ਤੋਂ ਮਹੱਤਵਪੂਰਨ cost ਲੁਕਾਉਂਦਾ ਹੈ।

ਪਹਿਲਾਂ ਮੁਫ਼ਤ runtime settings

Hardware upgrade ਪਹਿਲਾ performance step ਨਹੀਂ ਹੈ। Shopping page ਖੋਲ੍ਹਣ ਤੋਂ ਪਹਿਲਾਂ runtime settings test ਕਰੋ।

Multi-Token Prediction

ਕੁਝ model ਅਤੇ backend combinations ਕਈ future tokens predict ਕਰਕੇ ਇੱਕ pass ਵਿੱਚ verify ਕਰਦੇ ਹਨ। ਇਹ feature ਅਕਸਰ runtime flag ਜਾਂ compatible draft setup ਰਾਹੀਂ ਮਿਲਦਾ ਹੈ।

Reference material ਵਿੱਚ report ਕੀਤੇ tests ਕੁਝ high-end cards ਉੱਤੇ ਵੱਡੇ gains ਦਿਖਾਉਂਦੇ ਹਨ। Results model file, backend, driver ਅਤੇ card ਅਨੁਸਾਰ ਬਦਲਦੇ ਹਨ। Apple Metal paths conversion ਦੌਰਾਨ ਲੋੜੀਂਦਾ feature ਕਾਇਮ ਨਹੀਂ ਰੱਖ ਸਕਦੇ।

Reasoning Level

Maximum reasoning enabled ਵਾਲਾ model result ਵਾਪਸ ਕਰਨ ਤੋਂ ਪਹਿਲਾਂ internal work ਉੱਤੇ ਵੱਧ ਸਮਾਂ ਲਗਾਉਂਦਾ ਹੈ। Medium reasoning code repair ਲਈ ਅਕਸਰ ਵਧੀਆ balance ਦਿੰਦੀ ਹੈ, ਖਾਸ ਕਰਕੇ ਜਦੋਂ task ਵਿੱਚ ਸਪਸ਼ਟ error message ਅਤੇ narrow file target ਹੋਵੇ।

ਇੱਕ ਸਧਾਰਨ test matrix ਵਰਤੋ:

  1. ਇੱਕੋ bug-fix task low, medium ਅਤੇ high reasoning ਨਾਲ ਚਲਾਓ।
  2. Time to first token, total time, patch success ਅਤੇ test result ਲਿਖੋ।
  3. Short prompt ਅਤੇ repository-sized prompt ਨਾਲ ਦੁਹਰਾਓ।
  4. ਸਭ ਤੋਂ ਵਧੀਆ completed task ਦੇਣ ਵਾਲੀ setting ਰੱਖੋ, highest token rate ਨਹੀਂ।

Compatible multi-token path ਨਾਲ medium reasoning ਇੱਕ ਮਜ਼ਬੂਤ starting point ਹੈ। ਇਸਨੂੰ default ਬਣਾਉਣ ਤੋਂ ਪਹਿਲਾਂ ਆਪਣੇ code ਉੱਤੇ quality verify ਕਰੋ।

Local Hardware ਜਾਂ Rental GPU?

Occasional work ਲਈ rental compute ਜਿੱਤਦਾ ਹੈ। ਤੁਸੀਂ ਸਾਲ ਭਰ card ਖਰੀਦਣ, cool ਕਰਨ, update ਕਰਨ ਅਤੇ power ਦੇਣ ਦੀ ਬਜਾਏ active sessions ਲਈ pay ਕਰਦੇ ਹੋ।

Workload frequent, private ਜਾਂ offline ਹੋਣ ਉੱਤੇ owned hardware ਜਿੱਤਦਾ ਹੈ। ਇਹ queue time ਹਟਾਉਂਦਾ ਹੈ ਅਤੇ repeatable tests ਲਈ stable environment ਦਿੰਦਾ ਹੈ।

SituationBetter first move
A few sessions per monthGPU rent ਕਰੋ ਜਾਂ API ਵਰਤੋ
Daily private codingSupported 24GB ਤੋਂ 32GB system ਖਰੀਦੋ
Large repository with frequent reloadsਪਹਿਲਾਂ rent ਕਰੋ ਅਤੇ prefill speed ਮਾਪੋ
No code leaves the buildingContext target ਪੂਰਾ ਕਰਨ ਵਾਲਾ ਸਭ ਤੋਂ ਛੋਟਾ system own ਕਰੋ
Experimenting with a new modelHardware ਖਰੀਦਣ ਤੋਂ ਪਹਿਲਾਂ rent ਕਰੋ

Break-even ਨੂੰ calendar hours ਨਾਲ ਨਹੀਂ, active hours ਨਾਲ calculate ਕਰੋ। Electricity, storage, cooling, maintenance ਅਤੇ runtime ਚਲਦਾ ਰੱਖਣ ਲਈ ਲੱਗਣ ਵਾਲਾ ਸਮਾਂ ਸ਼ਾਮਲ ਕਰੋ।

Occasional experimentation ਲਈ ਖਰੀਦਿਆ high-end GPU hobby expense ਹੈ। Private work ਲਈ ਰੋਜ਼ ਵਰਤਿਆ 24GB ਤੋਂ 32GB system ਵਧੀਆ economic case ਰੱਖਦਾ ਹੈ।

ਬਿਹਤਰ Buying Checklist

Model ਅਤੇ GPU ਦੀ ਤੁਲਨਾ ਕਰਦੇ ਸਮੇਂ ਇਹ ਕ੍ਰਮ ਵਰਤੋ:

  1. Task define ਕਰੋ। Autocomplete, single-file repair, repository agent ਜਾਂ long-context analysis।
  2. Prompt ਮਾਪੋ। Normal system instructions, tool schemas, files ਅਤੇ test output ਗਿਣੋ।
  3. Cache behavior inspect ਕਰੋ। Architecture notes ਅਤੇ context-memory measurements ਵੇਖੋ।
  4. Quantization ਚੁਣੋ। ਸਿਰਫ਼ bit count ਨਹੀਂ, ਖਾਸ upload ਦੇ quality results ਜਾਂਚੋ।
  5. Prefill ਅਤੇ decode test ਕਰੋ। ਆਪਣੇ repository ਦੇ prompts ਵਰਤੋ।
  6. Reasoning tune ਕਰੋ। ਕਈ reasoning levels ਉੱਤੇ completed task time compare ਕਰੋ।
  7. Privacy ਅਤੇ maintenance test ਕਰੋ। Confirm ਕਰੋ ਕਿ source code ਕਿੱਥੇ ਜਾਂਦਾ ਹੈ ਅਤੇ backend ਕੌਣ maintain ਕਰਦਾ ਹੈ।
  8. Rental cost compare ਕਰੋ। Expected active hours ਵਰਤੋ ਅਤੇ owned-system estimate ਵਿੱਚ power ਸ਼ਾਮਲ ਕਰੋ।

ਅੰਤਿਮ ਸਿਫਾਰਸ਼

12GB ਤੋਂ ਘੱਟ ਉੱਤੇ ਵੱਧ system RAM ਵਾਲਾ ਛੋਟਾ model ਜਾਂ mixture-of-experts model ਚਲਾਓ। 27B dense model ਨੂੰ ਅਜਿਹੇ setup ਵਿੱਚ ਨਾ ਥੋਪੋ ਜੋ ਜ਼ਿਆਦਾਤਰ ਸਮਾਂ offload ਕਰਦਾ ਰਹੇ।

12GB ਤੋਂ 16GB ਵਿੱਚ quantization quality ਅਤੇ controlled context target ਉੱਤੇ ਧਿਆਨ ਦਿਓ। Tool support ਵਾਲਾ well-tested 2-bit ਜਾਂ 3-bit build ਉਸ 4-bit file ਨਾਲੋਂ ਵੱਧ ਲਾਭਦਾਇਕ ਹੈ ਜੋ cleanly fit ਨਹੀਂ ਹੁੰਦਾ।

24GB ਤੋਂ 32GB ਵਿੱਚ 27B coding model private daily work ਲਈ practical default ਬਣ ਜਾਂਦਾ ਹੈ। Full advertised window ਉਪਲਬਧ ਮੰਨਣ ਤੋਂ ਪਹਿਲਾਂ context usage ਅਤੇ prompt speed ਮਾਪੋ।

48GB ਅਤੇ ਵੱਧ ਉੱਤੇ ਵਾਧੂ memory precision, cache room ਅਤੇ stable sessions ਲਈ ਵਰਤੋ, ਵੱਡੇ model ਵੱਲ ਜਾਣ ਤੋਂ ਪਹਿਲਾਂ। ਜਦੋਂ model 128GB territory ਵਿੱਚ ਪਹੁੰਚਦਾ ਹੈ, prompt-reading speed ਅਤੇ rental economics raw capacity ਨਾਲੋਂ ਵੱਧ ਧਿਆਨ ਮੰਗਦੇ ਹਨ।

Model name ਸਿਰਫ਼ starting point ਹੈ। Useful question ਇਹ ਹੈ ਕਿ model, runtime, tools ਅਤੇ cache ਆਪਣਾ ਹਿੱਸਾ ਲੈਣ ਤੋਂ ਬਾਅਦ project context ਕਿੰਨਾ ਬਚਦਾ ਹੈ।

ਸੰਬੰਧਿਤ ਪਾਠ