# my-deepseek: offline DeepSeek-R1 chat server on RHEL 9.6, used from Windows VS Code (Continue)

| Item | Value |
|---|---|
| Model | **DeepSeek-R1-Distill-Qwen-14B**, Q4_K_M, Ollama tag `deepseek-r1:14b` (digest `c333b7232bdb`, 9.0 GB, MIT license) |
| Server engine | Ollama **0.34.1**, official Linux x86-64 build (includes CPU, CUDA 12/13 and Vulkan backends) |
| Address | `127.0.0.1:11550` by default (set in `deepseek.env`) |
| API | Ollama native `/api/chat` (used by Continue) and OpenAI-compatible `/v1/chat/completions` |
| VS Code extension | Continue **2.0.0**, `continue-win32-x64-2.0.0.vsix` (latest stable Marketplace release as of 2026-09-17) |

Everything needed offline is already in this directory. No internet is needed on the RHEL server or the Windows PC.

```
my-deepseek/
├── guide.md                 this file
├── deepseek.env             server settings (address, context size, ...)
├── start.sh stop.sh status.sh test.sh verify.sh
├── download.sh              ONLINE only: re-download the model (already done)
├── my-deepseek.service      optional systemd unit (auto-start at boot)
├── continue-config.yaml     Continue config for Windows
├── runtime/                 Ollama binary and libraries (extracted, ready to run)
├── models/                  deepseek-r1:14b model (manifest + blobs)
├── MODEL-LICENSE.txt
└── downloads/
    ├── continue-win32-x64-2.0.0.vsix   (+ .sha256)   → copy to the Windows PC
    └── ollama-linux-amd64.tar.zst      (+ sha256sum.txt) original Ollama archive
```

---

## 1. Requirements

**RHEL 9.6 server (x86-64)**

- RAM: at least **16 GB**. The model uses about 10 GB with an 8,192-token context.
- Disk: 12 GB for the bundle without `downloads/`, or 14 GB with it.
- Packages (all in a standard RHEL 9 install): `bash`, `curl`, `python3`, `tar`, `coreutils`, `util-linux`, `procps-ng`.
- GPU (optional): with an NVIDIA GPU that has **≥ 12 GB VRAM**, the whole model runs on the GPU and is much faster. The NVIDIA driver is **not** bundled. Install it on the RHEL server before you disconnect it from the network. Without a GPU, the model runs on the CPU.

**Windows PC**

- VS Code 1.70 or newer (64-bit x64). Use an up-to-date VS Code.
- OpenSSH client (built into Windows 10 and 11) if you use the SSH-tunnel connection (section 5).

---

## 2. Copy the bundle to the offline RHEL server

On the machine that has this directory:

```bash
cd /work/acb
tar --exclude=my-deepseek/logs/* --exclude=my-deepseek/run/* \
    -cf my-deepseek-offline.tar my-deepseek
sha256sum my-deepseek-offline.tar > my-deepseek-offline.tar.sha256
```

(Add `--exclude=my-deepseek/downloads` to leave out the 1.5 GB of installers. Then copy the `.vsix` to the Windows PC separately.)

Transfer the `.tar` and `.sha256` files (USB drive or internal file share). Then, on the RHEL server:

```bash
sha256sum -c my-deepseek-offline.tar.sha256
sudo tar -xf my-deepseek-offline.tar -C /opt       # → /opt/my-deepseek
cd /opt/my-deepseek
./verify.sh                                       # checks SHA-256 of every model file (~1 min)
```

`verify.sh` must end with every line `OK` and exit code 0. You can place the bundle in any directory, because the scripts find their own location. `/opt/my-deepseek` is only needed if you use the systemd unit as it is.

---

## 3. Start, test and stop the server

```bash
cd /opt/my-deepseek
./start.sh      # starts in the background, never downloads anything
./test.sh       # asks "What is 17 * 23?" and prints the thinking + answer
./status.sh     # shows the installed model and CPU/GPU placement
./stop.sh
```

- The first request loads 9 GB into memory, which can take a while. Later requests are faster. The model stays loaded for 30 minutes after the last request.
- You can use your own prompt: `./test.sh "Explain what a Linux cgroup is in 3 sentences."`
- To chat in the terminal:
  ```bash
  OLLAMA_HOST=127.0.0.1:11550 ./runtime/bin/ollama run deepseek-r1:14b
  ```
- Log: `logs/server.log`. The last test response: `logs/test-response.json`.
- R1 always "thinks" first (a reasoning section) and then answers. `test.sh` prints both sections.

### Settings (`deepseek.env`)

| Variable | Default | Notes |
|---|---|---|
| `OLLAMA_HOST` | `127.0.0.1:11550` | `0.0.0.0:11550` makes the server reachable from the LAN (see section 5B) |
| `OLLAMA_CONTEXT_LENGTH` | `8192` | Reasoning uses many tokens, so don't go below 8192. Raising it uses more RAM/VRAM. |
| `OLLAMA_NUM_PARALLEL` | `1` | Number of requests handled at the same time |
| `OLLAMA_KEEP_ALIVE` | `30m` | How long the model stays in memory after the last request |
| `OLLAMA_NO_CLOUD` | `1` | Turns off Ollama cloud features. This does not block the network. |

Restart after editing: `./stop.sh && ./start.sh`.

### Optional: start automatically at boot (systemd)

The unit file expects the bundle at `/opt/my-deepseek`. Don't run it together with `start.sh`.

```bash
cd /opt/my-deepseek && ./stop.sh
sudo cp my-deepseek.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now my-deepseek
systemctl status my-deepseek
journalctl -u my-deepseek -f          # log
```

If SELinux is `Enforcing` and the service fails with a permission error, check `sudo ausearch -m avc -ts recent`. Running `sudo restorecon -R /opt/my-deepseek` usually fixes the file labels.

---

## 4. Install Continue on the Windows PC (offline)

1. Copy `downloads/continue-win32-x64-2.0.0.vsix` to the Windows PC.
2. Optional: check the file in PowerShell:
   ```powershell
   Get-FileHash .\continue-win32-x64-2.0.0.vsix -Algorithm SHA256
   ```
   Expected hash: `D681ACA70F24E1A3B9ADEA8D9BBE4D24D954AE74F71C7B648ED589E53DC8354B`
3. Install the extension in one of two ways:
   - In VS Code, open Extensions (`Ctrl+Shift+X`), click `…` (top right), choose **Install from VSIX…**, and select the file.
   - Or in a terminal: `code --install-extension continue-win32-x64-2.0.0.vsix`
4. Reload VS Code. The Continue icon appears in the left activity bar.

This VSIX is for **Windows x64**. For Windows on ARM, download `continue-win32-arm64` instead.

---

## 5. Connect Windows to the RHEL server

Continue on Windows must be able to reach port 11550 on the RHEL server. Use **one** of these methods.

### A. SSH tunnel (recommended; leave the server on 127.0.0.1)

Ollama has **no authentication**. A tunnel keeps the server closed to the network. Keep this PowerShell window open while you use Continue:

```powershell
ssh -N -L 11550:127.0.0.1:11550 youruser@RHEL-SERVER-IP
```

Continue's `apiBase` is then `http://127.0.0.1:11550`, which is already set in the supplied config.

### B. Direct LAN access

On the RHEL server:

```bash
# in deepseek.env
OLLAMA_HOST=0.0.0.0:11550
```
```bash
./stop.sh && ./start.sh      # or: sudo systemctl restart my-deepseek
# allow only your Windows PCs' subnet (example 192.168.10.0/24)
sudo firewall-cmd --permanent --add-rich-rule='rule family="ipv4" source address="192.168.10.0/24" port port="11550" protocol="tcp" accept'
sudo firewall-cmd --reload
```

In the Continue config, set `apiBase: http://RHEL-SERVER-IP:11550`.

### Check the connection from Windows

```powershell
Invoke-RestMethod http://127.0.0.1:11550/api/tags        # tunnel
Invoke-RestMethod http://RHEL-SERVER-IP:11550/api/tags   # direct LAN
```

The output must list `deepseek-r1:14b`.

---

## 6. Configure Continue

Continue reads its configuration from **`%USERPROFILE%\.continue\config.yaml`** (for example `C:\Users\you\.continue\config.yaml`).

1. Open the Continue panel. If it shows a sign-in or onboarding screen, skip or close it. No account is needed for a local model.
2. Open the config: click the settings/gear icon in the Continue panel and open the **local config**. Or open the file directly in VS Code:
   ```powershell
   code $env:USERPROFILE\.continue\config.yaml
   ```
   (If the folder or file doesn't exist yet, create it.)
3. Back up any existing content. Then replace the file with this (the same content as `continue-config.yaml` in this bundle):

```yaml
name: Local DeepSeek R1
version: 1.0.0
schema: v1
models:
  - name: DeepSeek-R1-Distill-Qwen-14B
    provider: ollama
    model: deepseek-r1:14b
    # SSH tunnel (section 5A):  http://127.0.0.1:11550
    # Direct LAN (section 5B):  http://RHEL-SERVER-IP:11550
    apiBase: http://127.0.0.1:11550
    roles:
      - chat
      - edit
      - apply
    defaultCompletionOptions:
      contextLength: 8192
      maxTokens: 4096
      temperature: 0.6
      topP: 0.95
```

What each setting does:

| Key | Why |
|---|---|
| `provider: ollama` | Uses Ollama's native `/api/chat`. Continue 2.0.0 displays R1's reasoning as a separate collapsible "thinking" block. |
| `model: deepseek-r1:14b` | Must exactly match the name shown by `./status.sh` / `/api/tags`. |
| `apiBase` | Server URL **without** `/v1`. |
| `roles: chat, edit, apply` | This model is used for Chat and for code edits. **No `autocomplete`**: a reasoning model is too slow for tab completion and isn't trained for it. |
| `contextLength: 8192` | Keep equal to `OLLAMA_CONTEXT_LENGTH` in `deepseek.env`. |
| `maxTokens: 4096` | Room for the reasoning plus the answer. Too low cuts answers off in the middle of the thinking. |
| `temperature 0.6`, `topP 0.95` | Sampling values recommended by DeepSeek for R1 models. Lower temperatures cause repetition. |

4. Save the file. Continue reloads it automatically. If it doesn't, run **Developer: Reload Window**.

### Recommended VS Code settings for offline use

Open **Preferences: Open User Settings (JSON)** and add:

```json
{
  "continue.telemetryEnabled": false,
  "continue.enableTabAutocomplete": false,
  "continue.enableNextEdit": false,
  "extensions.autoUpdate": false,
  "extensions.autoCheckUpdates": false,
  "telemetry.telemetryLevel": "off"
}
```

Autocomplete and Next Edit are turned off because no autocomplete model is configured.

---

## 7. Chat

1. Open the Continue panel (`Ctrl+L`, or the Continue icon).
2. In the model selector, choose **DeepSeek-R1-Distill-Qwen-14B**.
3. Use **Chat** mode. Ask, for example: `Write a Python function is_prime(n) with a short explanation.`
4. The reasoning appears first as a "thinking" block, followed by the answer.
5. **Edit:** select some code in the editor, press `Ctrl+I`, and describe the change (for example "add type hints"). Review the diff, then accept or reject it.
6. Add context with `@` (for example `@Files`, `@Current File`) or by selecting code before pressing `Ctrl+L`.

Tips for R1:
- Put all instructions in the user message. DeepSeek recommends no custom system prompt.
- Use **Chat** mode, not **Agent** mode. Ollama lists this model with a "tools" capability, but the R1 distill is not reliable at tool calls or autonomous file editing.
- Speed depends on hardware. The measured speed on the preparation host (CPU only, no GPU) was about **4 tokens/s**. A short question took about 55 s, and a small coding task with long reasoning (~1,100 tokens) took about 4.5 min. A GPU with ≥ 12 GB VRAM is many times faster.

---

## 8. Troubleshooting

| Symptom | Fix |
|---|---|
| `start.sh`: *Model deepseek-r1:14b not found* | `models/` is missing or incomplete. Copy it again and run `./verify.sh`. |
| `start.sh`: *Port 11550 is used by another server* | Change the port in `OLLAMA_HOST` **and** in Continue's `apiBase`. |
| Continue: *Failed to fetch* / connection refused | Check that the server is running (`./status.sh`), that the SSH tunnel window is still open, the firewall rule, and the `Invoke-RestMethod` test in section 5. |
| Continue: *model not found* | The `model:` value must be exactly `deepseek-r1:14b`. |
| Answer stops in the middle of the thinking | Increase `maxTokens` (for example 6000) and, if needed, the context length on both sides (`deepseek.env` and the config). |
| Very slow, `status.sh` shows `100% CPU` although there is a GPU | Check that the NVIDIA driver is installed (`nvidia-smi`) and look at `logs/server.log` for CUDA errors. If the GPU has too little VRAM, the model is split between CPU and GPU. |
| Out of memory | Close other large processes, or lower `OLLAMA_CONTEXT_LENGTH` to 4096 (and `contextLength` in the config; set `maxTokens` to 2048). |
| Systemd service fails with `$HOME is not defined` | Keep the `Environment=HOME=...` line in the unit file. |

---

## 9. Re-downloading (only on a machine with internet)

The model is already downloaded. To rebuild the bundle from scratch:

```bash
# Ollama runtime
curl -LO https://github.com/ollama/ollama/releases/download/v0.34.1/ollama-linux-amd64.tar.zst
tar --zstd -xf ollama-linux-amd64.tar.zst -C runtime
# Model (resumable)
./download.sh
# Continue VSIX (Windows x64). The Marketplace sends it gzip-compressed:
curl -L "https://marketplace.visualstudio.com/_apis/public/gallery/publishers/Continue/vsextensions/continue/2.0.0/vspackage?targetPlatform=win32-x64" \
  | gunzip > continue-win32-x64-2.0.0.vsix
```

---

## Validation performed (2026-09-17, on the preparation host)

- The Ollama archive SHA-256 matched the release's `sha256sum.txt`. All 5 model blobs matched their SHA-256 names (`verify.sh`). The VSIX is a valid archive with no errors: publisher `Continue`, version 2.0.0, target `win32-x64`, not a pre-release.
- `ollama show`: architecture qwen2, 14.8B parameters, Q4_K_M, MIT license.
- The server was started with **no network access** (network namespace, `unshare -n`; DNS lookup failed as expected). Results:
  - `test.sh` answered correctly through `/v1/chat/completions`, including reasoning.
  - A coding prompt produced a correct `is_prime`.
  - A streaming `/api/chat` request with `think: true` (the request format Continue 2.0.0 sends) returned separate thinking and content at 4.08 tok/s.
- `status.sh`: 10 GB loaded, context 8192, 100% CPU (the host has no GPU).
- The systemd unit passed `systemd-analyze verify`, and a trial run served `/api/tags`. The `HOME` fix was found and applied during this test.
- **Not tested here:** the Continue UI on a real Windows PC, GPU acceleration, and RHEL 9.6 itself (the preparation host runs AlmaLinux 9.8, which is binary-compatible). Do the checks in sections 2, 3, 5 and 7 on the target machines.
