1. SIMY
  2. Guides
  3. Run Qwen on your own servers

A guide to building a local, in-house LLM

Handle confidential talk with your own Qwen, safely.

Customer data and design discussions can't be pasted into an outside AI. So run Qwen on your own servers, in a way where nothing leaves. Here are the steps and a checklist, based on the official documentation.

  • Pick from Apache 2.0 models
  • Start on 127.0.0.1, block outbound traffic
  • Includes a 20-item checklist

* This setup is an illustration

Contents
  1. Sound familiar?
  2. What you really wanted
  3. Local LLMs and why Qwen
  4. Choosing an inference engine
  5. Hardware guidelines
  6. Setup in a closed network
  7. Security checklist
  8. Common mistakes
  9. What the videos say
  10. Where SIMY fits
  11. FAQ
  12. Sources
01 / CAN'T SEND IT OUT

Confidential data you can't paste into an outside AI. Sound familiar?

You know AI would be faster. But these conversations can't leave the company.

e.g. a file with personal data The call came 3 minutes later

Someone had a personal AI account anonymize it. Infosec called, and a report had to be filed.

e.g. documents with customer data The project stalled for 6 months

Legal said "it can't go to an outside AI." Everyone sees the value, yet nothing moves.

A design review meeting It works, but it's worrying

You got it running in Ollama. But you can't say who is able to connect to it.

It can't go out, so you don't use AI. Between those two options lies "run it on your own servers."

* 1 and 2 are examples based on cases described in explainer videos (What the videos say).

02 / WHAT YOU REALLY WANTED

What you really wanted: confidential data, with Qwen in-house

Same documents, same meetings. AI does the work, and the data stays inside the company.

Still confidential AI summaries and drafts

Data stays on your own servers. No need to paste it into an outside AI.

When IT asks Explain it on one page

Where it listens, authentication, outbound traffic, logs. Show where each is locked down.

When upgrading Update without breaking

Fetch, verify, record, deploy. A fixed flow that works the same whoever runs it.

What comes back to you Only the decision

"Do we move to this version?" The before-and-after comparison is ready; that's the one question.

You can build all of this with the steps and checklist on this page. How to split the work, including the follow-up after meetings, is covered in Where SIMY fits.

03 / BASICS

What is a local LLM? Why run Qwen in-house

A local LLM means running the published weights of a model on your own PC or your company's servers. The text you enter never leaves that machine.

NameWhere it runsGood for
Local LLMYour own PC (Mac or Windows)Trying it alone, using it offline
On-premises / in-house LLMCompany servers or data centerShared use across a department or the company
Closed-network LLMInside a network not connected to the internetConfidentiality duties, personal data, design data

What a local LLM can do

  • Summarize and draft confidential documents: meeting minutes, contracts, design documents, customer correspondence.
  • Search and ask questions over internal documents (RAG): feed it internal policies and manuals and have it answer questions.
  • Translation and code help: translation between Chinese, English and other languages, explaining internal code and suggesting fixes.

Small models are not as smart as large cloud models. Use them to produce first drafts rather than final versions, and expectations will match reality more closely.

Why choose Qwen

  • Apache 2.0 models
  • Many languages
  • A choice of sizes
  • License: many models, such as Qwen3 and Qwen3.8-27B, are released under Apache 2.0. But not all of them (see the table below).
  • Languages: with the Qwen3 generation, Qwen claims support for 119 languages and dialects.
  • Size: from 0.6B to 8B that run on a PC, to 27B to 32B that fit on one GPU when quantized, up to 235B that needs multiple GPUs.

Licenses differ by model

Scroll the table sideways →

ModelLicenseHow confirmed
Qwen3 (0.6B to 32B, 30B-A3B, 235B-A22B-2507), Qwen3-VL, Qwen3-Coder, Qwen3-Embedding and RerankerApache 2.0Checked on Hugging Face
Qwen3.5 (0.8B to 35B-A3B), Qwen3.6 (27B, 35B-A3B), Qwen3.8-27B (including the FP8 version)Apache 2.0Checked on Hugging Face
Most of Qwen2.5 (0.5B to 32B, Coder 7B and 14B, VL-7B)Apache 2.0Checked on Hugging Face
Qwen2.5-3B, Qwen2.5-VL-3Bqwen-research (for research; commercial use handled separately)Name only. Check the full terms
Qwen2.5-72B-Instruct, Qwen2.5-VL-72Bqwen (custom license)Name only. Check the conditions
Qwen3.8-2.4T-A95BCustom licenseChecked on Hugging Face
Qwen3.8-Flash-Nextqwen-community-1.0. Offering it as MaaS, or using it as an "AI work assistant" for coding or office support, reportedly requires a separate licenseConfirmed from a summary. Legal review required

Checked against the model information in the Qwen organization on Hugging Face on October 1, 2026. Not every Qwen model is Apache 2.0. Before commercial use, check the LICENSE in the repository of the model you use with your legal team.

Recommended local LLMs: choosing a Qwen model by use caseAs of October 2026. When in doubt, start small
  • Try it on a PC first: Qwen3 4B to 8B, or the small Qwen3.5 models. They run in Ollama or LM Studio, so you can check whether they help your work.
  • Summaries and drafts for a department: Qwen3-14B or 32B, or Qwen3.8-27B. FP8 or 4-bit versions fit on a single GPU, and 3.8-27B also handles images.
  • Code help: Qwen3-Coder (30B-A3B). It is an MoE with few active parameters, so it runs fast for its size.
  • Searching internal documents: combine the answering model with Qwen3-Embedding and the Reranker.

Decide on one use before you test, and it becomes easier to judge which size is enough. Before buying a large GPU, the quickest path is to check on the PC you already have whether it works as a first-draft tool.

How to use Qwen itself (Qwen Chat, the API, the main models) is covered in the Qwen Guide, and comparisons with other Chinese models in Chinese AI models (LLMs) compared.

04 / ENGINES

Choosing an inference engine: Ollama, llama.cpp, vLLM, SGLang, LM Studio

The software that runs a model is the "inference engine." For security, the biggest difference is where it listens if you change nothing.

Scroll the table sideways →

EngineDefault listen addressBuilt-in authTLSOutbound trafficGood for
Ollama127.0.0.1:11434Not covered in the official FAQ. Handle it in a proxyNone. Terminate at a proxyCloud features stop with OLLAMA_NO_CLOUD=1. The Mac and Windows apps fetch updates automaticallyTrials by individuals or small teams, Mac
llama.cpp (llama-server)127.0.0.1:8080--api-key, --api-key-file--ssl-key-file, --ssl-cert-file--offline stops network checksSingle user, GGUF, CPU, Apple, AMD
vLLMAll interfaces if you omit --host (port 8000)--api-key or VLLM_API_KEY. Only some endpoints such as /v1 are protected--ssl-keyfile, --ssl-certfileSends usage stats by default. Stop with VLLM_NO_USAGE_STATS=1Production with many concurrent users across a department or company
SGLang127.0.0.1:30000--api-key. --admin-api-key for adminSafer to terminate at a proxyStop model downloads with HF_HUB_OFFLINE=1Agents, reusing the same long prompt
LM StudioYour own PC only (port 1234). LAN exposure is a settingCheck the official docsNoneCheck the official docsTrying it in a GUI, Macs and personal PCs

Checked against each project's official documentation (and the vLLM source) as of October 2026. In the LM Studio row, items we could not fully confirm are marked "Check the official docs."

  • Check first: use llama.cpp or Ollama to see whether it helps your work.
  • Company-wide use: vLLM. Speed holds up even with many concurrent users.
  • Agents sending the same long instructions in volume: SGLang.

A vLLM caveat: the official Security page explains that --api-key only protects endpoints such as /v1, and not /invocations, /pooling and others. That is why the project itself names a reverse proxy that passes only the endpoints you want to expose as the most effective measure.

05 / HARDWARE

Hardware guidelines: GPU memory needed for each Qwen model

Think of the GPU memory you need as the sum of three parts.

  • WeightsParameters × bytes per parameter (BF16 is 2, FP8 is 1, 4-bit is about 0.5)
  • + KV cacheProportional to context length × concurrent users
  • + HeadroomWorking space at run time

Scroll the table sideways →

ModelBF16FP84-bit (INT4 etc.)Source
Qwen3-8BAbout 16 GBAbout 9.3 GBAbout 6.2 GBQwen's official speed benchmark
Qwen3-14BAbout 28 GBAbout 16 GBAbout 10 GBSame
Qwen3-32BAbout 63 GBAbout 33 GBAbout 19 GBSame
Qwen3-235B-A22B8 GPUs4 GPUs4 GPUs (GPTQ-INT4)Same (SGLang setup)
Qwen3.8-27BAbout 54 GB (weights only, calculated estimate)About 27 GB (calculated). One measurement puts the official FP8 version at about 28.8 GBAbout 14 GB (calculated). Actual files are around 17 to 18 GB (the Ollama download is about 18 GB)Calculated from parameter count, a measurement example, the Ollama library

The Qwen3 figures come from Qwen's official speed benchmark (GPU memory right after startup, processing a short input with transformers). Qwen3.8-27B has no official table, so its figures are estimates calculated from the parameter count plus an individual's measurements, and they exclude the KV cache. The amount needed varies widely with the quantization and context length (as of October 2026).

  • Qwen3.8-27B and a 32 GB GPU: BF16 needs about 54 GB for the weights alone (a calculated estimate), so it doesn't fit on one RTX 5090 (32 GB). The official FP8 version is also about 28.8 GB for the weights alone, and one estimate shows it won't fit in 32 GB at the maximum context. Adjust by using a 4-bit version or lowering the context limit.
  • Two 16 GB GPUs are not the same as one 32 GB GPU: splitting across two cards costs you the inter-card traffic and the working space each card needs.
  • Long contexts mean longer waits: in one individual's measurement, an input of about 250,000 tokens took over 4 minutes before the first character appeared.
  • Mac: Apple silicon shares memory with the GPU. You need memory well above the size of the model file.

How to size it: decide first how many people will use it at once and how long the documents they feed it will be. The KV cache grows in proportion to both, so if you assume the weights alone are enough, you'll run short in production.

06 / BUILD

Setup in a closed network: Qwen + vLLM on an internal server

The example uses the FP8 version of Qwen3.8-27B with vLLM (Docker) on a Linux server. Commands match the official documentation as of October 2026. If GPU memory runs short, lower the context limit (--max-model-len) or choose a 4-bit version.

* This setup is an illustration

Reading the commands: replace anything in angle brackets, such as <verified commit SHA>, with your own values. Version numbers are examples as of October 2026.

  1. Fetch the model at a pinned version

    Transfer machine (internet side)

    Download from the official Qwen/ organization's repository, specifying a commit. Save the inference engine's container at a pinned version too.

    bash
    pip install -U huggingface_hub
    hf download Qwen/Qwen3.8-27B-FP8 \
      --revision <verified commit SHA> \
      --local-dir ./Qwen3.8-27B-FP8
    
    docker pull vllm/vllm-openai:v0.30.0
    docker inspect --format '{{index .RepoDigests 0}}' vllm/vllm-openai:v0.30.0
    docker save vllm/vllm-openai:v0.30.0 -o vllm-openai-v0.30.0.tar

    Copy the digest that docker inspect prints (the value starting with sha256:) into your ledger.

    Why pin the version: even in the same repository, contents can change with a new revision. In one quantized version, a new revision reportedly removed the layers used for speculative decoding (MTP), so the speed-up stopped working.

  2. Build a SHA-256 ledger

    Transfer machine

    The lfs.oid returned by the Hugging Face API is the SHA-256 of a large (LFS) file. Check your local files against these values.

    bash
    # Extract the values from Hugging Face (LFS files only)
    curl -s "https://huggingface.co/api/models/Qwen/Qwen3.8-27B-FP8/tree/<verified commit SHA>?recursive=true&expand=true" \
      | jq -r '.[] | select(.lfs) | "\(.lfs.oid)  ./\(.path)"' > HF.sha256
    
    # Check the local files and build a ledger of every file
    cd Qwen3.8-27B-FP8
    sha256sum -c ../HF.sha256
    find . -type f ! -path './.cache/*' -print0 | sort -z | xargs -0 sha256sum > ../MODEL.sha256
    cd .. && sha256sum vllm-openai-v0.30.0.tar > IMAGE.sha256

    Have someone other than the person who downloaded the files sign off on the ledger (follow your internal standard for signing).

  3. Bring it in on media and verify before loading

    Closed-network server

    Move the ledger and files on media, and check them again on the closed side. If even one line says FAILED, don't use it.

    bash
    cd /srv/llm/models/Qwen3.8-27B-FP8 && sha256sum -c /srv/llm/incoming/MODEL.sha256
    cd /srv/llm/incoming && sha256sum -c IMAGE.sha256
    docker load -i vllm-openai-v0.30.0.tar
  4. Start it listening on 127.0.0.1 only

    Closed-network server

    vLLM listens on all interfaces if you omit --host. Always write --host 127.0.0.1, and pass the key in a file (on the command line it is visible in ps).

    bash · vLLM (Docker)
    # Create the key file with permissions 600
    sudo install -D -m 600 /dev/null /etc/llm/vllm.env
    echo "VLLM_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/llm/vllm.env > /dev/null
    
    sudo docker run -d --name qwen --gpus all --ipc=host --network host \
      -v /srv/llm/models:/models:ro \
      -e HF_HUB_OFFLINE=1 -e VLLM_NO_USAGE_STATS=1 -e DO_NOT_TRACK=1 \
      --env-file /etc/llm/vllm.env \
      vllm/vllm-openai:v0.30.0 \
      --model /models/Qwen3.8-27B-FP8 \
      --served-model-name qwen3.8-27b \
      --host 127.0.0.1 --port 8000 \
      --max-model-len 65536 --gpu-memory-utilization 0.90
    
    # Confirm it listens on 127.0.0.1:8000 only
    ss -ltnp | grep 8000
    • Don't add --trust-remote-code. This keeps code bundled with the model from running.
    • Don't use VLLM_SERVER_DEV_MODE=1 in production. It exposes development endpoints.
    • Thinking mode and tool calls: Qwen's official documentation gives --reasoning-parser qwen3 and --enable-auto-tool-choice --tool-call-parser hermes as examples for Qwen3. Check the model card for the settings for the 3.8 generation.
    bash · With Ollama (systemd)
    sudo systemctl edit ollama
    # Add these 3 lines to the file that opens, then save
    [Service]
    Environment="OLLAMA_HOST=127.0.0.1:11434"
    Environment="OLLAMA_NO_CLOUD=1"
    
    sudo systemctl restart ollama
    bash · With SGLang or llama.cpp
    python3 -m sglang.launch_server --model-path /models/Qwen3.8-27B-FP8 \
      --host 127.0.0.1 --port 30000 --api-key <key> --admin-api-key <admin key>
    
    llama-server -m /models/qwen.gguf --host 127.0.0.1 --port 8080 \
      --api-key-file /etc/llm/keys --offline --no-webui --jinja

    In the SGLang example too, don't write the keys directly; in practice pass them from a file or environment variable. The model card examples use --host 0.0.0.0, so don't copy them as they are.

  5. Add TLS, an allowlist and authentication with nginx

    Closed-network server (gateway)

    Users touch only nginx. Pass only the parts of /v1 you need, and return 404 for everything else.

    nginx
    server {
      listen 443 ssl;
      server_name llm.internal.example;
      ssl_certificate     /etc/pki/llm.crt;
      ssl_certificate_key /etc/pki/llm.key;
      ssl_protocols TLSv1.2 TLSv1.3;
    
      allow 10.20.0.0/16;   # internal segment only
      deny  all;
    
      location = /v1/chat/completions {
        auth_request /_auth;
        auth_request_set $user $upstream_http_x_auth_request_user;
        proxy_set_header Authorization "Bearer <upstream key>";
        proxy_pass http://127.0.0.1:8000;
        proxy_buffering off;   # for streaming responses
      }
      location = /v1/models {
        auth_request /_auth;
        auth_request_set $user $upstream_http_x_auth_request_user;
        proxy_set_header Authorization "Bearer <upstream key>";
        proxy_pass http://127.0.0.1:8000;
      }
      location = /_auth {
        internal;
        proxy_pass http://127.0.0.1:4180/oauth2/auth;   # connect to your internal IdP via oauth2-proxy or similar
      }
      location / { return 404; }   # don't expose /metrics, /invocations and the like
    }

    Users sign in with your internal SSO (OIDC or SAML), and the upstream API key is never handed out to users. auth_request is an nginx module; check whether your build includes it with nginx -V.

  6. Block outbound traffic

    Closed-network server

    Block all traffic from the server to the internet by default. Allow only internal DNS, time sync and your log collector.

    bash · ufw
    sudo ufw default deny incoming
    sudo ufw default deny outgoing
    sudo ufw allow in on <admin IF> to any port 22 proto tcp    # allow this first or SSH drops
    sudo ufw allow in on <internal IF> to any port 443 proto tcp
    sudo ufw allow out to <internal DNS> port 53 proto udp
    sudo ufw allow out to <internal NTP> port 123 proto udp
    sudo ufw allow out to <log collector> port 514 proto tcp
    sudo ufw enable
    
    # Confirm it can't reach the outside (failure means it's correct)
    curl -m 5 https://huggingface.co

    If you publish a port with Docker's -p, Docker adds its own packet rules, so traffic from outside can get through even when you think ufw has closed it. These steps use --network host and --host 127.0.0.1 to keep listening local.

    bash · When using -p
    # Publish on 127.0.0.1 only on the host (-p 8000:8000 publishes on all interfaces)
    sudo docker run -d --name qwen --gpus all --ipc=host \
      -p 127.0.0.1:8000:8000 \
      -v /srv/llm/models:/models:ro \
      -e HF_HUB_OFFLINE=1 -e VLLM_NO_USAGE_STATS=1 -e DO_NOT_TRACK=1 \
      --env-file /etc/llm/vllm.env \
      vllm/vllm-openai:v0.30.0 \
      --model /models/Qwen3.8-27B-FP8 --served-model-name qwen3.8-27b \
      --host 0.0.0.0 --port 8000 --max-model-len 65536
    # ↑ listens on 0.0.0.0 inside the container, exposed on 127.0.0.1 only on the host

    In this form, vLLM inside the container listens on 0.0.0.0, but it can't be reached from outside the host. Block outbound traffic at the network firewall as well.

  7. Keep logs and audit trails

    Gateway and monitoring

    Record who used which endpoint and when in the nginx logs, and send them to your internal log platform (such as a SIEM). Count tokens with /metrics (Prometheus format), exposed only to the monitoring network.

    nginx · inside http { }
    log_format llm '$time_iso8601 user=$user ip=$remote_addr "$request" '
                   'status=$status bytes=$body_bytes_sent rt=$request_time';
    access_log /var/log/nginx/llm.log llm;

    If you enable --enable-log-requests in vLLM at DEBUG level, the prompt text ends up in the logs. Decide whether to keep the text, and for how long, according to your rules on personal and confidential data.

  8. Decide users and permissions

    UI (Open WebUI or similar)

    If people will chat through a UI, put a front end such as Open WebUI in the same closed network, and limit which models each role and group can use. Make sure people who leave are cut off through your internal IdP.

    env · Open WebUI
    OFFLINE_MODE=true
    ENABLE_SIGNUP=false
    DEFAULT_USER_ROLE=pending
    ENABLE_COMMUNITY_SHARING=false

    With OFFLINE_MODE enabled, document search (RAG) won't work unless the embedding model is installed beforehand. Bring in the embedding model through the same flow as steps 1 to 3.

  9. Set up the update flow in advance

    Transfer machine → test environment → production

    vLLM releases roughly every two weeks, and almost every release includes security fixes. Even in a closed network, upgrade regularly through "fetch → verify → record → production."

    bash
    # 1. Fetch, check and bring in the new version as in steps 1 to 3
    # 2. In the test environment, send the same questions to the old and new versions and compare
    # 3. List the environment variable names in use and check against the release notes that none were removed
    sudo docker inspect --format '{{range .Config.Env}}{{println .}}{{end}}' qwen \
      | grep -E '^(VLLM_|HF_)' | cut -d= -f1
    # 4. Once in production, log one line on who upgraded when, and when it was rolled back
    echo "$(date -I) vllm <old version>→<new version> owner:<name> verified:OK rollback:<old version>.tar" >> /srv/llm/CHANGELOG

    Keep the previous container and model, with their digests noted, so you can roll back. Compare values that change between versions, such as the context length limit, before and after.

07 / CHECKLIST

Local LLM security checklist: 20 items before production

Go through these from top to bottom before and after setup. The checkboxes work only on this screen and nothing is sent anywhere.

0 / 20 checked

Getting the model and containers

Listening and outbound traffic

Entry point and authentication

Permissions, records and operations

08 / MISTAKES

Common mistakes: the gaps that "I thought it was fine" leaves in a local LLM

Common mistakes on the left, the fix from this page's steps on the right.

  • --host 0.0.0.0 with no authUsing a sample start command as is and exposing it to the whole LAN.
    127.0.0.1 + nginxListen locally only. Users go through the gateway.
  • Modified versions of unknown originUsing community GGUFs such as "Uncensored" builds without checking where they came from.
    From the official Qwen/, pinned by SHAIf you use a third-party version, record its author and process.
  • "It has an api-key, so it's safe"Assuming vLLM's key protects every endpoint.
    Narrow it with an allowlistPass only what you need, such as /v1/chat/completions.
  • "It's on-premises, so it's safe"Skipping permissions, logs and RAG access control because it sits inside the company.
    Manage it in-house like anything externalSSO, roles, logs, and search that respects access rights.
  • Using it from outside via a tunnelExposing an internal tool to the outside with Cloudflare Tunnel or similar.
    Only from the internal networkIf outside access is needed, go through the company's official remote access.
  • Missing default outbound trafficvLLM's usage stats, or auto-updates in the Ollama desktop app.
    Both environment variables and outbound blockingStop it in settings, and stop it at the firewall too.
  • Trusting small models too muchIn agents and tool calls, they invent files or links that don't exist.
    Use it for first draftsRun things in a sandbox, and have a person check the results.
  • Breaking on updateAfter upgrading, removed environment variables or changed limits stop it from running.
    Compare before and after in a test environmentCheck against the release notes and record who upgraded it.
09 / FROM THE VIDEOS

What the videos say: real examples of self-hosting Qwen

We went through 18 explainer videos in English, Chinese and Japanese published between February and October 2026, down to their audio transcripts.

What the 18 videos say8 scenes, and what they show as a whole

From what was said: 8 scenes

  1. What happened at a bank (secondhand): a speaker describes a case at a client's bank: an employee had a personally subscribed AI tool read a file containing personal data to anonymize it, and three minutes later the infosec team called; it ended in a report and disciplinary action (paraphrased).
    Will 保哥 (talk) · from 1:39:00 in the video
  2. The harness is the risk: in the same talk, the speaker says the danger in agents is not the model but the harness (the machinery around it that writes files and talks to the network); the model is just stringing characters together inside the GPU (paraphrased). This is the reason for item 19 (sandbox) on the checklist.
    Will 保哥 (talk) · from 1:21:00 in the video
  3. Why pin the version: the creator says a new revision of a quantized version removed the MTP layers, so speculative decoding stopped working, which is why you pin to a commit (paraphrased). This is why step 1 says to pin the version.
    RepoChad · from 4:05 in the video
  4. Port 11434 has no lock: the creator says Ollama's API on localhost:11434 has no authentication by default, so on a server you control it with a reverse proxy or firewall (paraphrased).
    Satsuki's OSS Lab (in Japanese) · from 10:00 in the video
  5. Don't open port 8000 to the world: in an example on a cloud GPU, the creator says they don't want port 8000 open to the world, so they connect through SSH port forwarding (paraphrased). It stands in for TLS during testing, before TLS is set up.
    The Cef Experience · from 9:31 in the video
  6. Too visible, even on-premises: the creator says that if you neglect permission design, the AI will answer with salary and HR information people should never see; putting it in-house alone does not make it safe (paraphrased).
    Kanri no Pro (in Japanese) · from 24:30 in the video
  7. One line per update: the creator says to pin the current version and keep the tag, and log one line on who upgraded when and when it was rolled back, because it pays off during incidents (paraphrased). This is the record format in step 9.
    AI News (in Japanese) · from 7:01 in the video
  8. Time is a cost too: after two days of trying Ollama with a tunnel and then going back to the cloud, the creator says "free" has hidden costs, and time is a cost too (paraphrased).
    Garfield_Investment · from 8:02 in the video

Some videos demonstrate using unofficial modified models to bypass the licensing of commercial software. Don't copy them. Modified versions of unknown origin are not used in this page's steps.

What the 18 videos show as a whole

  • Confidentiality duties: a tax accountant built a tax AI with one Mac, Ollama, Qwen3 and Open WebUI that keeps working with Wi-Fi turned off. He also says multi-step tax calculations are hard, and that for now the best option is to use the cloud with the client's consent.
  • Air gap: four stages: mirror the tools, physically bring in the model, and start the inference engine from storage inside the closed network. Evaluation and a gateway are part of the flow too.
  • Choosing an engine: llama.cpp for one person, vLLM for many, SGLang for agents that reuse the same long instructions. Some also argue for Linux over Windows in production.
  • Two GPUs don't become one big GPU: several explained that inter-card traffic becomes the bottleneck, so speed for a single user barely changes; what you gain is concurrent connections and context length.
  • Waiting on long contexts: in an example running Qwen3.8-27B on two 16 GB GPUs, an input of about 250,000 tokens took over 4 minutes before the first character.

Statements are paraphrased from automatic transcripts of the audio, not verbatim quotes. Timestamps are approximate; the point may come about 30 seconds after the link. Figures are one example for each setup.

10 / SIMY

Where SIMY fits: Qwen inside the closed network, SIMY outside

To be upfront: SIMY does not currently run on Qwen, and it does not run inside a closed network or offline.

* This division of work is an illustration

  • Picks up from meetings and chat
  • Your way of working
  • Returns only decisions

Confidential content is handled by Qwen in the closed network. After that, for "who does what by when," SIMY moves forward only the part the company has decided may go outside, within tools the company allows.

  • How SIMY runs: cloud runs use your own ChatGPT account (Codex). Codex usage falls within your ChatGPT plan.
  • Running on your PC: the SIMY desktop app can use Claude Code on your PC to do the work. Download it here.
  • Reading chat history: the desktop app reads your Claude Code and Codex chat history and turns work you repeat into "Suggestions from SIMY." Chat content is not stored on our servers. Decide which PCs and what scope according to your company's rules.

Where to draw the line: your company's rules decide which information stays in the closed network and what may be handled by outside tools. How SIMY handles data is published in Security and the Privacy Policy.

11 / FAQ

Local LLM and Qwen FAQ

What is a local LLM?

It means running the published weights of a model on your own PC or on your company's servers. The text you enter never leaves that machine, so you can use it with confidential documents.

What can a local LLM do?

It can summarize and draft confidential documents, answer questions over your internal documents (RAG), and help with translation and code. Small models are better at producing first drafts than final versions.

Which local LLM do you recommend?

To try things out, Qwen3 4B to 8B; for a department, Qwen3-14B or 32B, or Qwen3.8-27B; for code, Qwen3-Coder (as of October 2026). Before buying a big GPU, check on the PC you already have whether it works as a first-draft tool.

Is Qwen free if I run it locally?

Apache 2.0 models have no usage fees and can be used commercially. But the GPUs, servers, electricity and operations are on you, and some models use their own licenses.

What specs do I need to run Qwen locally?

According to Qwen's official figures, Qwen3-8B uses about 16 GB of GPU memory in BF16 and Qwen3-32B about 33 GB in FP8. These figures are mostly the weights; the KV cache grows on top of that in proportion to the number of concurrent users and the context length.

Does Qwen3.8-27B run on a single RTX 5090?

A 4-bit quantized version fits on one 32 GB RTX 5090. BF16 needs about 54 GB for the weights alone (a calculated estimate), so it does not fit, and the FP8 version is about 29 GB for the weights alone, so it only fits if you cut the context limit substantially.

What is the difference between Ollama and vLLM?

Ollama is for trying things out easily; vLLM is an engine for production use by many people at once. Ollama listens only on 127.0.0.1 by default, but vLLM listens on all interfaces if you omit --host, so always set it.

Do I still need updates in a closed network?

Yes. vLLM releases roughly every two weeks, and almost every release includes security fixes. Set up a flow where you fetch updates on a transfer machine and compare them in a test environment before putting them into production.

Can Qwen be used commercially?

Apache 2.0 models (such as Qwen3 and Qwen3.8-27B) can be used commercially. Qwen2.5-3B, the 72B models, Qwen3.8-Flash-Next and others use different licenses, so check the LICENSE of the model you use with your legal team.

Which languages does Qwen support?

Many. With the Qwen3 generation, Qwen claims support for 119 languages and dialects. Small models can sound unnatural in long texts, so check important wording before you use it.

Does Qwen run on a Mac?

Yes. Ollama, LM Studio and llama.cpp all support the Mac. Apple silicon shares memory with the GPU, so you need memory well above the size of the model file.

Is on-premises enough to be safe?

Not on its own. Only once you have decided where it listens, authentication, outbound traffic, logging and access rights to internal documents can you explain that nothing leaves.

Can SIMY be used in a closed network?

No. At this time SIMY does not run inside a closed network or offline, and it does not run on Qwen. You can split the work: handle confidential content with Qwen in the closed network, and use SIMY for follow-up work within the scope your company allows.

12 / SOURCES

Sources

Steps, default settings and licenses were checked against the following official sources (checked October 1, 2026). Versions and defaults change often, so check each official page before you build.

Confidential work stays with your own Qwen.
The follow-up goes to SIMY.

Inside the closed network, use the steps and checklist on this page. For work within the scope your company allows, SIMY picks it up from meetings and chat and moves it forward, returning only the decisions to you.