Contents
Confidential data you can't paste into an outside AI. Sound familiar?
You know AI would be faster. But these conversations can't leave the company.
Someone had a personal AI account anonymize it. Infosec called, and a report had to be filed.
Legal said "it can't go to an outside AI." Everyone sees the value, yet nothing moves.
You got it running in Ollama. But you can't say who is able to connect to it.
It can't go out, so you don't use AI. Between those two options lies "run it on your own servers."
* 1 and 2 are examples based on cases described in explainer videos (What the videos say).
What you really wanted: confidential data, with Qwen in-house
Same documents, same meetings. AI does the work, and the data stays inside the company.
Data stays on your own servers. No need to paste it into an outside AI.
Where it listens, authentication, outbound traffic, logs. Show where each is locked down.
Fetch, verify, record, deploy. A fixed flow that works the same whoever runs it.
"Do we move to this version?" The before-and-after comparison is ready; that's the one question.
You can build all of this with the steps and checklist on this page. How to split the work, including the follow-up after meetings, is covered in Where SIMY fits.
What is a local LLM? Why run Qwen in-house
A local LLM means running the published weights of a model on your own PC or your company's servers. The text you enter never leaves that machine.
| Name | Where it runs | Good for |
|---|---|---|
| Local LLM | Your own PC (Mac or Windows) | Trying it alone, using it offline |
| On-premises / in-house LLM | Company servers or data center | Shared use across a department or the company |
| Closed-network LLM | Inside a network not connected to the internet | Confidentiality duties, personal data, design data |
What a local LLM can do
- Summarize and draft confidential documents: meeting minutes, contracts, design documents, customer correspondence.
- Search and ask questions over internal documents (RAG): feed it internal policies and manuals and have it answer questions.
- Translation and code help: translation between Chinese, English and other languages, explaining internal code and suggesting fixes.
Small models are not as smart as large cloud models. Use them to produce first drafts rather than final versions, and expectations will match reality more closely.
Why choose Qwen
- Apache 2.0 models
- Many languages
- A choice of sizes
- License: many models, such as Qwen3 and Qwen3.8-27B, are released under Apache 2.0. But not all of them (see the table below).
- Languages: with the Qwen3 generation, Qwen claims support for 119 languages and dialects.
- Size: from 0.6B to 8B that run on a PC, to 27B to 32B that fit on one GPU when quantized, up to 235B that needs multiple GPUs.
Licenses differ by model
Scroll the table sideways →
| Model | License | How confirmed |
|---|---|---|
| Qwen3 (0.6B to 32B, 30B-A3B, 235B-A22B-2507), Qwen3-VL, Qwen3-Coder, Qwen3-Embedding and Reranker | Apache 2.0 | Checked on Hugging Face |
| Qwen3.5 (0.8B to 35B-A3B), Qwen3.6 (27B, 35B-A3B), Qwen3.8-27B (including the FP8 version) | Apache 2.0 | Checked on Hugging Face |
| Most of Qwen2.5 (0.5B to 32B, Coder 7B and 14B, VL-7B) | Apache 2.0 | Checked on Hugging Face |
| Qwen2.5-3B, Qwen2.5-VL-3B | qwen-research (for research; commercial use handled separately) | Name only. Check the full terms |
| Qwen2.5-72B-Instruct, Qwen2.5-VL-72B | qwen (custom license) | Name only. Check the conditions |
| Qwen3.8-2.4T-A95B | Custom license | Checked on Hugging Face |
| Qwen3.8-Flash-Next | qwen-community-1.0. Offering it as MaaS, or using it as an "AI work assistant" for coding or office support, reportedly requires a separate license | Confirmed from a summary. Legal review required |
Checked against the model information in the Qwen organization on Hugging Face on October 1, 2026. Not every Qwen model is Apache 2.0. Before commercial use, check the LICENSE in the repository of the model you use with your legal team.
Recommended local LLMs: choosing a Qwen model by use caseAs of October 2026. When in doubt, start small
- Try it on a PC first: Qwen3 4B to 8B, or the small Qwen3.5 models. They run in Ollama or LM Studio, so you can check whether they help your work.
- Summaries and drafts for a department: Qwen3-14B or 32B, or Qwen3.8-27B. FP8 or 4-bit versions fit on a single GPU, and 3.8-27B also handles images.
- Code help: Qwen3-Coder (30B-A3B). It is an MoE with few active parameters, so it runs fast for its size.
- Searching internal documents: combine the answering model with Qwen3-Embedding and the Reranker.
Decide on one use before you test, and it becomes easier to judge which size is enough. Before buying a large GPU, the quickest path is to check on the PC you already have whether it works as a first-draft tool.
How to use Qwen itself (Qwen Chat, the API, the main models) is covered in the Qwen Guide, and comparisons with other Chinese models in Chinese AI models (LLMs) compared.
Choosing an inference engine: Ollama, llama.cpp, vLLM, SGLang, LM Studio
The software that runs a model is the "inference engine." For security, the biggest difference is where it listens if you change nothing.
Scroll the table sideways →
| Engine | Default listen address | Built-in auth | TLS | Outbound traffic | Good for |
|---|---|---|---|---|---|
| Ollama | 127.0.0.1:11434 | Not covered in the official FAQ. Handle it in a proxy | None. Terminate at a proxy | Cloud features stop with OLLAMA_NO_CLOUD=1. The Mac and Windows apps fetch updates automatically | Trials by individuals or small teams, Mac |
| llama.cpp (llama-server) | 127.0.0.1:8080 | --api-key, --api-key-file | --ssl-key-file, --ssl-cert-file | --offline stops network checks | Single user, GGUF, CPU, Apple, AMD |
| vLLM | All interfaces if you omit --host (port 8000) | --api-key or VLLM_API_KEY. Only some endpoints such as /v1 are protected | --ssl-keyfile, --ssl-certfile | Sends usage stats by default. Stop with VLLM_NO_USAGE_STATS=1 | Production with many concurrent users across a department or company |
| SGLang | 127.0.0.1:30000 | --api-key. --admin-api-key for admin | Safer to terminate at a proxy | Stop model downloads with HF_HUB_OFFLINE=1 | Agents, reusing the same long prompt |
| LM Studio | Your own PC only (port 1234). LAN exposure is a setting | Check the official docs | None | Check the official docs | Trying it in a GUI, Macs and personal PCs |
Checked against each project's official documentation (and the vLLM source) as of October 2026. In the LM Studio row, items we could not fully confirm are marked "Check the official docs."
- Check first: use llama.cpp or Ollama to see whether it helps your work.
- Company-wide use: vLLM. Speed holds up even with many concurrent users.
- Agents sending the same long instructions in volume: SGLang.
A vLLM caveat: the official Security page explains that --api-key only protects endpoints such as /v1, and not /invocations, /pooling and others. That is why the project itself names a reverse proxy that passes only the endpoints you want to expose as the most effective measure.
Hardware guidelines: GPU memory needed for each Qwen model
Think of the GPU memory you need as the sum of three parts.
- WeightsParameters × bytes per parameter (BF16 is 2, FP8 is 1, 4-bit is about 0.5)
- + KV cacheProportional to context length × concurrent users
- + HeadroomWorking space at run time
Scroll the table sideways →
| Model | BF16 | FP8 | 4-bit (INT4 etc.) | Source |
|---|---|---|---|---|
| Qwen3-8B | About 16 GB | About 9.3 GB | About 6.2 GB | Qwen's official speed benchmark |
| Qwen3-14B | About 28 GB | About 16 GB | About 10 GB | Same |
| Qwen3-32B | About 63 GB | About 33 GB | About 19 GB | Same |
| Qwen3-235B-A22B | 8 GPUs | 4 GPUs | 4 GPUs (GPTQ-INT4) | Same (SGLang setup) |
| Qwen3.8-27B | About 54 GB (weights only, calculated estimate) | About 27 GB (calculated). One measurement puts the official FP8 version at about 28.8 GB | About 14 GB (calculated). Actual files are around 17 to 18 GB (the Ollama download is about 18 GB) | Calculated from parameter count, a measurement example, the Ollama library |
The Qwen3 figures come from Qwen's official speed benchmark (GPU memory right after startup, processing a short input with transformers). Qwen3.8-27B has no official table, so its figures are estimates calculated from the parameter count plus an individual's measurements, and they exclude the KV cache. The amount needed varies widely with the quantization and context length (as of October 2026).
- Qwen3.8-27B and a 32 GB GPU: BF16 needs about 54 GB for the weights alone (a calculated estimate), so it doesn't fit on one RTX 5090 (32 GB). The official FP8 version is also about 28.8 GB for the weights alone, and one estimate shows it won't fit in 32 GB at the maximum context. Adjust by using a 4-bit version or lowering the context limit.
- Two 16 GB GPUs are not the same as one 32 GB GPU: splitting across two cards costs you the inter-card traffic and the working space each card needs.
- Long contexts mean longer waits: in one individual's measurement, an input of about 250,000 tokens took over 4 minutes before the first character appeared.
- Mac: Apple silicon shares memory with the GPU. You need memory well above the size of the model file.
How to size it: decide first how many people will use it at once and how long the documents they feed it will be. The KV cache grows in proportion to both, so if you assume the weights alone are enough, you'll run short in production.
Setup in a closed network: Qwen + vLLM on an internal server
The example uses the FP8 version of Qwen3.8-27B with vLLM (Docker) on a Linux server. Commands match the official documentation as of October 2026. If GPU memory runs short, lower the context limit (--max-model-len) or choose a 4-bit version.
* This setup is an illustration
Reading the commands: replace anything in angle brackets, such as <verified commit SHA>, with your own values. Version numbers are examples as of October 2026.
-
Fetch the model at a pinned version
Transfer machine (internet side)Download from the official
Qwen/organization's repository, specifying a commit. Save the inference engine's container at a pinned version too.bashpip install -U huggingface_hub hf download Qwen/Qwen3.8-27B-FP8 \ --revision <verified commit SHA> \ --local-dir ./Qwen3.8-27B-FP8 docker pull vllm/vllm-openai:v0.30.0 docker inspect --format '{{index .RepoDigests 0}}' vllm/vllm-openai:v0.30.0 docker save vllm/vllm-openai:v0.30.0 -o vllm-openai-v0.30.0.tarCopy the digest that
docker inspectprints (the value starting withsha256:) into your ledger.Why pin the version: even in the same repository, contents can change with a new revision. In one quantized version, a new revision reportedly removed the layers used for speculative decoding (MTP), so the speed-up stopped working.
-
Build a SHA-256 ledger
Transfer machineThe
lfs.oidreturned by the Hugging Face API is the SHA-256 of a large (LFS) file. Check your local files against these values.bash# Extract the values from Hugging Face (LFS files only) curl -s "https://huggingface.co/api/models/Qwen/Qwen3.8-27B-FP8/tree/<verified commit SHA>?recursive=true&expand=true" \ | jq -r '.[] | select(.lfs) | "\(.lfs.oid) ./\(.path)"' > HF.sha256 # Check the local files and build a ledger of every file cd Qwen3.8-27B-FP8 sha256sum -c ../HF.sha256 find . -type f ! -path './.cache/*' -print0 | sort -z | xargs -0 sha256sum > ../MODEL.sha256 cd .. && sha256sum vllm-openai-v0.30.0.tar > IMAGE.sha256Have someone other than the person who downloaded the files sign off on the ledger (follow your internal standard for signing).
-
Bring it in on media and verify before loading
Closed-network serverMove the ledger and files on media, and check them again on the closed side. If even one line says
FAILED, don't use it.bashcd /srv/llm/models/Qwen3.8-27B-FP8 && sha256sum -c /srv/llm/incoming/MODEL.sha256 cd /srv/llm/incoming && sha256sum -c IMAGE.sha256 docker load -i vllm-openai-v0.30.0.tar -
Start it listening on 127.0.0.1 only
Closed-network servervLLM listens on all interfaces if you omit
--host. Always write--host 127.0.0.1, and pass the key in a file (on the command line it is visible inps).bash · vLLM (Docker)# Create the key file with permissions 600 sudo install -D -m 600 /dev/null /etc/llm/vllm.env echo "VLLM_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/llm/vllm.env > /dev/null sudo docker run -d --name qwen --gpus all --ipc=host --network host \ -v /srv/llm/models:/models:ro \ -e HF_HUB_OFFLINE=1 -e VLLM_NO_USAGE_STATS=1 -e DO_NOT_TRACK=1 \ --env-file /etc/llm/vllm.env \ vllm/vllm-openai:v0.30.0 \ --model /models/Qwen3.8-27B-FP8 \ --served-model-name qwen3.8-27b \ --host 127.0.0.1 --port 8000 \ --max-model-len 65536 --gpu-memory-utilization 0.90 # Confirm it listens on 127.0.0.1:8000 only ss -ltnp | grep 8000- Don't add
--trust-remote-code. This keeps code bundled with the model from running. - Don't use
VLLM_SERVER_DEV_MODE=1in production. It exposes development endpoints. - Thinking mode and tool calls: Qwen's official documentation gives
--reasoning-parser qwen3and--enable-auto-tool-choice --tool-call-parser hermesas examples for Qwen3. Check the model card for the settings for the 3.8 generation.
bash · With Ollama (systemd)sudo systemctl edit ollama # Add these 3 lines to the file that opens, then save [Service] Environment="OLLAMA_HOST=127.0.0.1:11434" Environment="OLLAMA_NO_CLOUD=1" sudo systemctl restart ollamabash · With SGLang or llama.cpppython3 -m sglang.launch_server --model-path /models/Qwen3.8-27B-FP8 \ --host 127.0.0.1 --port 30000 --api-key <key> --admin-api-key <admin key> llama-server -m /models/qwen.gguf --host 127.0.0.1 --port 8080 \ --api-key-file /etc/llm/keys --offline --no-webui --jinjaIn the SGLang example too, don't write the keys directly; in practice pass them from a file or environment variable. The model card examples use
--host 0.0.0.0, so don't copy them as they are. - Don't add
-
Add TLS, an allowlist and authentication with nginx
Closed-network server (gateway)Users touch only nginx. Pass only the parts of
/v1you need, and return 404 for everything else.nginxserver { listen 443 ssl; server_name llm.internal.example; ssl_certificate /etc/pki/llm.crt; ssl_certificate_key /etc/pki/llm.key; ssl_protocols TLSv1.2 TLSv1.3; allow 10.20.0.0/16; # internal segment only deny all; location = /v1/chat/completions { auth_request /_auth; auth_request_set $user $upstream_http_x_auth_request_user; proxy_set_header Authorization "Bearer <upstream key>"; proxy_pass http://127.0.0.1:8000; proxy_buffering off; # for streaming responses } location = /v1/models { auth_request /_auth; auth_request_set $user $upstream_http_x_auth_request_user; proxy_set_header Authorization "Bearer <upstream key>"; proxy_pass http://127.0.0.1:8000; } location = /_auth { internal; proxy_pass http://127.0.0.1:4180/oauth2/auth; # connect to your internal IdP via oauth2-proxy or similar } location / { return 404; } # don't expose /metrics, /invocations and the like }Users sign in with your internal SSO (OIDC or SAML), and the upstream API key is never handed out to users.
auth_requestis an nginx module; check whether your build includes it withnginx -V. -
Block outbound traffic
Closed-network serverBlock all traffic from the server to the internet by default. Allow only internal DNS, time sync and your log collector.
bash · ufwsudo ufw default deny incoming sudo ufw default deny outgoing sudo ufw allow in on <admin IF> to any port 22 proto tcp # allow this first or SSH drops sudo ufw allow in on <internal IF> to any port 443 proto tcp sudo ufw allow out to <internal DNS> port 53 proto udp sudo ufw allow out to <internal NTP> port 123 proto udp sudo ufw allow out to <log collector> port 514 proto tcp sudo ufw enable # Confirm it can't reach the outside (failure means it's correct) curl -m 5 https://huggingface.coIf you publish a port with Docker's
-p, Docker adds its own packet rules, so traffic from outside can get through even when you think ufw has closed it. These steps use--network hostand--host 127.0.0.1to keep listening local.bash · When using -p# Publish on 127.0.0.1 only on the host (-p 8000:8000 publishes on all interfaces) sudo docker run -d --name qwen --gpus all --ipc=host \ -p 127.0.0.1:8000:8000 \ -v /srv/llm/models:/models:ro \ -e HF_HUB_OFFLINE=1 -e VLLM_NO_USAGE_STATS=1 -e DO_NOT_TRACK=1 \ --env-file /etc/llm/vllm.env \ vllm/vllm-openai:v0.30.0 \ --model /models/Qwen3.8-27B-FP8 --served-model-name qwen3.8-27b \ --host 0.0.0.0 --port 8000 --max-model-len 65536 # ↑ listens on 0.0.0.0 inside the container, exposed on 127.0.0.1 only on the hostIn this form, vLLM inside the container listens on
0.0.0.0, but it can't be reached from outside the host. Block outbound traffic at the network firewall as well. -
Keep logs and audit trails
Gateway and monitoringRecord who used which endpoint and when in the nginx logs, and send them to your internal log platform (such as a SIEM). Count tokens with
/metrics(Prometheus format), exposed only to the monitoring network.nginx · inside http { }log_format llm '$time_iso8601 user=$user ip=$remote_addr "$request" ' 'status=$status bytes=$body_bytes_sent rt=$request_time'; access_log /var/log/nginx/llm.log llm;If you enable
--enable-log-requestsin vLLM at DEBUG level, the prompt text ends up in the logs. Decide whether to keep the text, and for how long, according to your rules on personal and confidential data. -
Decide users and permissions
UI (Open WebUI or similar)If people will chat through a UI, put a front end such as Open WebUI in the same closed network, and limit which models each role and group can use. Make sure people who leave are cut off through your internal IdP.
env · Open WebUIOFFLINE_MODE=true ENABLE_SIGNUP=false DEFAULT_USER_ROLE=pending ENABLE_COMMUNITY_SHARING=falseWith
OFFLINE_MODEenabled, document search (RAG) won't work unless the embedding model is installed beforehand. Bring in the embedding model through the same flow as steps 1 to 3. -
Set up the update flow in advance
Transfer machine → test environment → productionvLLM releases roughly every two weeks, and almost every release includes security fixes. Even in a closed network, upgrade regularly through "fetch → verify → record → production."
bash# 1. Fetch, check and bring in the new version as in steps 1 to 3 # 2. In the test environment, send the same questions to the old and new versions and compare # 3. List the environment variable names in use and check against the release notes that none were removed sudo docker inspect --format '{{range .Config.Env}}{{println .}}{{end}}' qwen \ | grep -E '^(VLLM_|HF_)' | cut -d= -f1 # 4. Once in production, log one line on who upgraded when, and when it was rolled back echo "$(date -I) vllm <old version>→<new version> owner:<name> verified:OK rollback:<old version>.tar" >> /srv/llm/CHANGELOGKeep the previous container and model, with their digests noted, so you can roll back. Compare values that change between versions, such as the context length limit, before and after.
Local LLM security checklist: 20 items before production
Go through these from top to bottom before and after setup. The checkboxes work only on this screen and nothing is sent anywhere.
0 / 20 checked
Getting the model and containers
Listening and outbound traffic
Entry point and authentication
Permissions, records and operations
Common mistakes: the gaps that "I thought it was fine" leaves in a local LLM
Common mistakes on the left, the fix from this page's steps on the right.
--host 0.0.0.0with no authUsing a sample start command as is and exposing it to the whole LAN.127.0.0.1+ nginxListen locally only. Users go through the gateway.- Modified versions of unknown originUsing community GGUFs such as "Uncensored" builds without checking where they came from.From the official
Qwen/, pinned by SHAIf you use a third-party version, record its author and process. - "It has an api-key, so it's safe"Assuming vLLM's key protects every endpoint.Narrow it with an allowlistPass only what you need, such as
/v1/chat/completions. - "It's on-premises, so it's safe"Skipping permissions, logs and RAG access control because it sits inside the company.Manage it in-house like anything externalSSO, roles, logs, and search that respects access rights.
- Using it from outside via a tunnelExposing an internal tool to the outside with Cloudflare Tunnel or similar.Only from the internal networkIf outside access is needed, go through the company's official remote access.
- Missing default outbound trafficvLLM's usage stats, or auto-updates in the Ollama desktop app.Both environment variables and outbound blockingStop it in settings, and stop it at the firewall too.
- Trusting small models too muchIn agents and tool calls, they invent files or links that don't exist.Use it for first draftsRun things in a sandbox, and have a person check the results.
- Breaking on updateAfter upgrading, removed environment variables or changed limits stop it from running.Compare before and after in a test environmentCheck against the release notes and record who upgraded it.
What the videos say: real examples of self-hosting Qwen
We went through 18 explainer videos in English, Chinese and Japanese published between February and October 2026, down to their audio transcripts.
What the 18 videos say8 scenes, and what they show as a whole
From what was said: 8 scenes
- What happened at a bank (secondhand): a speaker describes a case at a client's bank: an employee had a personally subscribed AI tool read a file containing personal data to anonymize it, and three minutes later the infosec team called; it ended in a report and disciplinary action (paraphrased).
Will 保哥 (talk) · from 1:39:00 in the video - The harness is the risk: in the same talk, the speaker says the danger in agents is not the model but the harness (the machinery around it that writes files and talks to the network); the model is just stringing characters together inside the GPU (paraphrased). This is the reason for item 19 (sandbox) on the checklist.
Will 保哥 (talk) · from 1:21:00 in the video - Why pin the version: the creator says a new revision of a quantized version removed the MTP layers, so speculative decoding stopped working, which is why you pin to a commit (paraphrased). This is why step 1 says to pin the version.
RepoChad · from 4:05 in the video - Port 11434 has no lock: the creator says Ollama's API on
localhost:11434has no authentication by default, so on a server you control it with a reverse proxy or firewall (paraphrased).
Satsuki's OSS Lab (in Japanese) · from 10:00 in the video - Don't open port 8000 to the world: in an example on a cloud GPU, the creator says they don't want port 8000 open to the world, so they connect through SSH port forwarding (paraphrased). It stands in for TLS during testing, before TLS is set up.
The Cef Experience · from 9:31 in the video - Too visible, even on-premises: the creator says that if you neglect permission design, the AI will answer with salary and HR information people should never see; putting it in-house alone does not make it safe (paraphrased).
Kanri no Pro (in Japanese) · from 24:30 in the video - One line per update: the creator says to pin the current version and keep the tag, and log one line on who upgraded when and when it was rolled back, because it pays off during incidents (paraphrased). This is the record format in step 9.
AI News (in Japanese) · from 7:01 in the video - Time is a cost too: after two days of trying Ollama with a tunnel and then going back to the cloud, the creator says "free" has hidden costs, and time is a cost too (paraphrased).
Garfield_Investment · from 8:02 in the video
Some videos demonstrate using unofficial modified models to bypass the licensing of commercial software. Don't copy them. Modified versions of unknown origin are not used in this page's steps.
What the 18 videos show as a whole
- Confidentiality duties: a tax accountant built a tax AI with one Mac, Ollama, Qwen3 and Open WebUI that keeps working with Wi-Fi turned off. He also says multi-step tax calculations are hard, and that for now the best option is to use the cloud with the client's consent.
- Air gap: four stages: mirror the tools, physically bring in the model, and start the inference engine from storage inside the closed network. Evaluation and a gateway are part of the flow too.
- Choosing an engine: llama.cpp for one person, vLLM for many, SGLang for agents that reuse the same long instructions. Some also argue for Linux over Windows in production.
- Two GPUs don't become one big GPU: several explained that inter-card traffic becomes the bottleneck, so speed for a single user barely changes; what you gain is concurrent connections and context length.
- Waiting on long contexts: in an example running Qwen3.8-27B on two 16 GB GPUs, an input of about 250,000 tokens took over 4 minutes before the first character.
Statements are paraphrased from automatic transcripts of the audio, not verbatim quotes. Timestamps are approximate; the point may come about 30 seconds after the link. Figures are one example for each setup.
Where SIMY fits: Qwen inside the closed network, SIMY outside
To be upfront: SIMY does not currently run on Qwen, and it does not run inside a closed network or offline.
- TranscriptConfidential meeting
- SummaryOn your own GPU
- DraftContent stays inside
- Allowed scopeOwners, deadlines, dates
- Owners and deadlinesPicked up from meetings
- ReplyDrafted in Gmail
- DecisionJust one question back
* This division of work is an illustration
- Picks up from meetings and chat
- Your way of working
- Returns only decisions
Confidential content is handled by Qwen in the closed network. After that, for "who does what by when," SIMY moves forward only the part the company has decided may go outside, within tools the company allows.
- How SIMY runs: cloud runs use your own ChatGPT account (Codex). Codex usage falls within your ChatGPT plan.
- Running on your PC: the SIMY desktop app can use Claude Code on your PC to do the work. Download it here.
- Reading chat history: the desktop app reads your Claude Code and Codex chat history and turns work you repeat into "Suggestions from SIMY." Chat content is not stored on our servers. Decide which PCs and what scope according to your company's rules.
Where to draw the line: your company's rules decide which information stays in the closed network and what may be handled by outside tools. How SIMY handles data is published in Security and the Privacy Policy.
Local LLM and Qwen FAQ
What is a local LLM?
It means running the published weights of a model on your own PC or on your company's servers. The text you enter never leaves that machine, so you can use it with confidential documents.
What can a local LLM do?
It can summarize and draft confidential documents, answer questions over your internal documents (RAG), and help with translation and code. Small models are better at producing first drafts than final versions.
Which local LLM do you recommend?
To try things out, Qwen3 4B to 8B; for a department, Qwen3-14B or 32B, or Qwen3.8-27B; for code, Qwen3-Coder (as of October 2026). Before buying a big GPU, check on the PC you already have whether it works as a first-draft tool.
Is Qwen free if I run it locally?
Apache 2.0 models have no usage fees and can be used commercially. But the GPUs, servers, electricity and operations are on you, and some models use their own licenses.
What specs do I need to run Qwen locally?
According to Qwen's official figures, Qwen3-8B uses about 16 GB of GPU memory in BF16 and Qwen3-32B about 33 GB in FP8. These figures are mostly the weights; the KV cache grows on top of that in proportion to the number of concurrent users and the context length.
Does Qwen3.8-27B run on a single RTX 5090?
A 4-bit quantized version fits on one 32 GB RTX 5090. BF16 needs about 54 GB for the weights alone (a calculated estimate), so it does not fit, and the FP8 version is about 29 GB for the weights alone, so it only fits if you cut the context limit substantially.
What is the difference between Ollama and vLLM?
Ollama is for trying things out easily; vLLM is an engine for production use by many people at once. Ollama listens only on 127.0.0.1 by default, but vLLM listens on all interfaces if you omit --host, so always set it.
Do I still need updates in a closed network?
Yes. vLLM releases roughly every two weeks, and almost every release includes security fixes. Set up a flow where you fetch updates on a transfer machine and compare them in a test environment before putting them into production.
Can Qwen be used commercially?
Apache 2.0 models (such as Qwen3 and Qwen3.8-27B) can be used commercially. Qwen2.5-3B, the 72B models, Qwen3.8-Flash-Next and others use different licenses, so check the LICENSE of the model you use with your legal team.
Which languages does Qwen support?
Many. With the Qwen3 generation, Qwen claims support for 119 languages and dialects. Small models can sound unnatural in long texts, so check important wording before you use it.
Does Qwen run on a Mac?
Yes. Ollama, LM Studio and llama.cpp all support the Mac. Apple silicon shares memory with the GPU, so you need memory well above the size of the model file.
Is on-premises enough to be safe?
Not on its own. Only once you have decided where it listens, authentication, outbound traffic, logging and access rights to internal documents can you explain that nothing leaves.
Can SIMY be used in a closed network?
No. At this time SIMY does not run inside a closed network or offline, and it does not run on Qwen. You can split the work: handle confidential content with Qwen in the closed network, and use SIMY for follow-up work within the scope your company allows.
Sources
Steps, default settings and licenses were checked against the following official sources (checked October 1, 2026). Versions and defaults change often, so check each official page before you build.
- Security (vLLM docs), Usage Stats Collection, vllm serve
- vLLM release notes (GitHub)
- Server Arguments (SGLang docs)
- FAQ (Ollama docs)
- llama.cpp server README (GitHub)
- Deploying with vLLM (Qwen docs), Speed Benchmark (Qwen docs)
- Qwen (Hugging Face), Qwen3.8-27B model card
- Hugging Face CLI, Environment variables (Hugging Face Hub)
- ngx_http_auth_request_module (nginx docs), ngx_http_ssl_module
- Environment Variable Configuration (Open WebUI docs)
- Packet filtering and firewalls (Docker docs)
- qwen3.8 (Ollama library)