惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

IT之家
IT之家
Y
Y Combinator Blog
月光博客
月光博客
Blog — PlanetScale
Blog — PlanetScale
GbyAI
GbyAI
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园 - 三生石上(FineUI控件)
S
SegmentFault 最新的问题
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
美团技术团队
雷峰网
雷峰网
酷 壳 – CoolShell
酷 壳 – CoolShell
Last Week in AI
Last Week in AI
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
有赞技术团队
有赞技术团队
博客园 - 司徒正美
V
Visual Studio Blog
小众软件
小众软件
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
T
Tailwind CSS Blog
Apple Machine Learning Research
Apple Machine Learning Research
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
A
About on SuperTechFans
The Cloudflare Blog

Proxmox Support Forum

[SOLVED] - Github Auth for Mirrors-Kernel Repo? [Automation] Mass migration tool for MS Win11/Server Proxmox GUI hang - not response is it possible to reject or quarantine spam based on conditions I set ? The PVENode task list in PVE9 is partially obscured due to the terminal font being too large. About 100% error reporting due to pveproxy.service hooks Kubernetes overlay networking breaks when upgrading from PVE 9.1 to PVE 9.2.3 Zentraler Speicher No space left on device Combine datastore and direct file archival to tape Kernel panic VFS: Unable to mount root fs on unknown-block (0,0) sobald ein 7.x Kernel verwendet wird. How to migrate disk of a VM from one ZFS to another Windows Server 2025 fails to boot after PVE 9.2 / Linux 7.0 Kernel upgrade Cannot Install Proxmox on T610 Poweredge with H700 PERC card sdn Config. gateway not reachable How to safely change domain/FQDN? Welche Filterquote erreicht ihr? NFS Share status unknown on 2 of 5 nodes Can't connect to PVE9 consoles [solved] Can't connect to PVE9 consoles [solved] [SOLVED] - Use secondary network for PVE commands Created cluster, one node storage gone BUG: proxmox mail gateway FROM = null bypass spam filtering Moving existing PBS from VMWare workstation to PVE cluster Does eBGP SDN fabric support external peering? Bug: PDM 1.1 not recognizing valid license status Proxmox GUI hang - not response Proxmox Backup Server 4.2 released! Advice ceph-osd crashes with kernel 6.17.2-1-pve on Dell system
PVE crashes unexpectedly
invalid@exam · 2026-05-30 · via Proxmox Support Forum

Hello everyone,
I could use some help to figure out why my PVE instance crashes randomly. I am running version 6.14.8-2-pve on a Lenovo ThinkStation P360 which I repurposed from a regular desktop machine to run PVE.

Pulling the logs, I can see the following when PVE becomes unavailable and I have to do a hard reboot:

Code:

May 15 00:00:00 pm kernel: e1000e 0000:00:1f.6 eno1: Detected Hardware Unit Hang:
                             TDH                  <b9>
                             TDT                  <d3>
                             next_to_use          <d3>
                             next_to_clean        <b8>
                           buffer_info[next_to_clean]:
                             time_stamp           <1257dfe98>
                             next_to_watch        <b9>
                             jiffies              <125ca9e00>
                             next_to_watch.status <0>
                           MAC Status             <80083>
                           PHY Status             <796d>
                           PHY 1000BASE-T Status  <3800>
                           PHY Extended Status    <3000>
                           PCI Status             <10>

Any suggestions what I can do to prevent PVE from crashing on me? I am using barely any hardware so it shouldn't be a heat related problem (I assume).

Last edited:

Thanks daanw, much appreciated! I ran the following command and will keep monitoring PVE for any potential hardware unit hangs.

Code:

ethtool -K eno1 gso off gro off tso off

The node has seen another crash within 48 hours. I'm adding the journalctl log from before the crash in case that helps figuring out the cause. Any suggestions?

Code:

May 31 10:51:14 pm pvedaemon[1282]: <root@pam> successful auth for user 'mirko@pve'
May 31 11:06:15 pm pvedaemon[1282]: <root@pam> successful auth for user 'mirko@pve'
May 31 11:17:01 pm CRON[238143]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 11:17:01 pm CRON[238145]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 11:17:01 pm CRON[238143]: pam_unix(cron:session): session closed for user root
May 31 11:21:16 pm pvedaemon[1280]: <root@pam> successful auth for user 'mirko@pve'
May 31 11:36:17 pm pvedaemon[1281]: <root@pam> successful auth for user 'mirko@pve'
May 31 11:42:52 pm pveproxy[123178]: worker exit
May 31 11:42:52 pm pveproxy[1291]: worker 123178 finished
May 31 11:42:52 pm pveproxy[1291]: starting 1 worker(s)
May 31 11:42:52 pm pveproxy[1291]: worker 242671 started
May 31 11:43:38 pm pveproxy[123177]: worker exit
May 31 11:43:38 pm pveproxy[1291]: worker 123177 finished
May 31 11:43:38 pm pveproxy[1291]: starting 1 worker(s)
May 31 11:43:38 pm pveproxy[1291]: worker 242809 started
May 31 11:44:12 pm pveproxy[123179]: worker exit
May 31 11:44:12 pm pveproxy[1291]: worker 123179 finished
May 31 11:44:12 pm pveproxy[1291]: starting 1 worker(s)
May 31 11:44:12 pm pveproxy[1291]: worker 242941 started
May 31 11:51:45 pm pvedaemon[1281]: <root@pam> successful auth for user 'mirko@pve'
May 31 12:07:45 pm pvedaemon[1280]: <root@pam> successful auth for user 'mirko@pve'
May 31 12:17:01 pm CRON[248560]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 12:17:01 pm CRON[248562]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 12:17:01 pm CRON[248560]: pam_unix(cron:session): session closed for user root
May 31 12:23:46 pm pvedaemon[1281]: <root@pam> successful auth for user 'mirko@pve'
May 31 12:49:15 pm systemd[1]: Starting systemd-tmpfiles-clean.service - Cleanup of Temporary Directories...
May 31 12:49:15 pm systemd-tmpfiles[254243]: /usr/lib/tmpfiles.d/legacy.conf:14: Duplicate line for path "/run/lock", ignoring.
May 31 12:49:15 pm systemd[1]: systemd-tmpfiles-clean.service: Deactivated successfully.
May 31 12:49:15 pm systemd[1]: Finished systemd-tmpfiles-clean.service - Cleanup of Temporary Directories.
May 31 13:17:01 pm CRON[259154]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 13:17:01 pm CRON[259156]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 13:17:01 pm CRON[259154]: pam_unix(cron:session): session closed for user root
May 31 14:17:01 pm CRON[269660]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 14:17:01 pm CRON[269662]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 14:17:01 pm CRON[269660]: pam_unix(cron:session): session closed for user root
May 31 14:38:20 pm kernel: hrtimer: interrupt took 13834 ns
May 31 15:17:01 pm CRON[280904]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 15:17:01 pm CRON[280906]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 15:17:01 pm CRON[280904]: pam_unix(cron:session): session closed for user root
May 31 15:20:13 pm systemd[1]: Starting apt-daily.service - Daily apt download activities...
May 31 15:20:13 pm systemd[1]: apt-daily.service: Deactivated successfully.
May 31 15:20:13 pm systemd[1]: Finished apt-daily.service - Daily apt download activities.
May 31 16:17:01 pm CRON[291806]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 16:17:01 pm CRON[291808]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 16:17:01 pm CRON[291806]: pam_unix(cron:session): session closed for user root
May 31 17:11:43 pm chronyd[1031]: Leap second list /usr/share/zoneinfo/leap-seconds.list needs update
May 31 17:17:01 pm CRON[302682]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 17:17:01 pm CRON[302684]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 17:17:01 pm CRON[302682]: pam_unix(cron:session): session closed for user root
May 31 18:17:01 pm CRON[313554]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 18:17:01 pm CRON[313556]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 18:17:01 pm CRON[313554]: pam_unix(cron:session): session closed for user root
May 31 19:17:01 pm CRON[324010]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 19:17:01 pm CRON[324012]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 19:17:01 pm CRON[324010]: pam_unix(cron:session): session closed for user root
May 31 20:17:01 pm CRON[334392]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 20:17:01 pm CRON[334394]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 20:17:01 pm CRON[334392]: pam_unix(cron:session): session closed for user root
May 31 21:10:13 pm systemd[1]: Starting apt-daily.service - Daily apt download activities...
May 31 21:10:13 pm systemd[1]: apt-daily.service: Deactivated successfully.
May 31 21:10:13 pm systemd[1]: Finished apt-daily.service - Daily apt download activities.
May 31 21:17:01 pm CRON[345399]: pam_unix(cron:session): session opened for user root(uid=0) by root(uid=0)
May 31 21:17:01 pm CRON[345401]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
May 31 21:17:01 pm CRON[345399]: pam_unix(cron:session): session closed for user root
May 31 21:56:21 pm kernel: e1000e 0000:00:1f.6 eno1: Detected Hardware Unit Hang:
                             TDH                  <f8>
                             TDT                  <aa>
                             next_to_use          <aa>
                             next_to_clean        <f7>
                           buffer_info[next_to_clean]:
                             time_stamp           <107257a8d>
                             next_to_watch        <f8>
                             jiffies              <107258240>
                             next_to_watch.status <0>
                           MAC Status             <80083>
                           PHY Status             <796d>
                           PHY 1000BASE-T Status  <3800>
                           PHY Extended Status    <3000>
                           PCI Status             <10>
May 31 21:56:23 pm kernel: e1000e 0000:00:1f.6 eno1: Detected Hardware Unit Hang:
                             TDH                  <f8>
                             TDT                  <aa>
                             next_to_use          <aa>
                             next_to_clean        <f7>
                           buffer_info[next_to_clean]:
                             time_stamp           <107257a8d>
                             next_to_watch        <f8>
                             jiffies              <107258a40>
                             next_to_watch.status <0>
                           MAC Status             <80083>
                           PHY Status             <796d>
                           PHY 1000BASE-T Status  <3800>
                           PHY Extended Status    <3000>
                           PCI Status             <10>

The node has seen another crash within 48 hours. Any suggestions? I'm not sure about the best way to share the log files, do let me know if those would help.

Code:

May 31 21:56:21 pm kernel: e1000e 0000:00:1f.6 eno1: Detected Hardware Unit Hang:
                             TDH                  <f8>
                             TDT                  <aa>
                             next_to_use          <aa>
                             next_to_clean        <f7>
                           buffer_info[next_to_clean]:
                             time_stamp           <107257a8d>
                             next_to_watch        <f8>
                             jiffies              <107258240>
                             next_to_watch.status <0>
                           MAC Status             <80083>
                           PHY Status             <796d>
                           PHY 1000BASE-T Status  <3800>
                           PHY Extended Status    <3000>
                           PCI Status             <10>
May 31 21:56:23 pm kernel: e1000e 0000:00:1f.6 eno1: Detected Hardware Unit Hang:
                             TDH                  <f8>
                             TDT                  <aa>
                             next_to_use          <aa>
                             next_to_clean        <f7>
                           buffer_info[next_to_clean]:
                             time_stamp           <107257a8d>
                             next_to_watch        <f8>
                             jiffies              <107258a40>
                             next_to_watch.status <0>
                           MAC Status             <80083>
                           PHY Status             <796d>
                           PHY 1000BASE-T Status  <3800>
                           PHY Extended Status    <3000>
                           PCI Status             <10>

Furthermore here is the config I set for the NIC:

Code:

root@pm:~# ethtool -k eno1
Features for eno1:
rx-checksumming: on
tx-checksumming: on
        tx-checksum-ipv4: off [fixed]
        tx-checksum-ip-generic: on
        tx-checksum-ipv6: off [fixed]
        tx-checksum-fcoe-crc: off [fixed]
        tx-checksum-sctp: off [fixed]
scatter-gather: on
        tx-scatter-gather: on
        tx-scatter-gather-fraglist: off [fixed]
tcp-segmentation-offload: off
        tx-tcp-segmentation: off
        tx-tcp-ecn-segmentation: off [fixed]
        tx-tcp-mangleid-segmentation: off
        tx-tcp6-segmentation: off
generic-segmentation-offload: off
generic-receive-offload: off
large-receive-offload: off [fixed]
rx-vlan-offload: on
tx-vlan-offload: on
ntuple-filters: off [fixed]
receive-hashing: on
highdma: on [fixed]
rx-vlan-filter: off [fixed]
vlan-challenged: off [fixed]
tx-gso-robust: off [fixed]
tx-fcoe-segmentation: off [fixed]
tx-gre-segmentation: off [fixed]
tx-gre-csum-segmentation: off [fixed]
tx-ipxip4-segmentation: off [fixed]
tx-ipxip6-segmentation: off [fixed]
tx-udp_tnl-segmentation: off [fixed]
tx-udp_tnl-csum-segmentation: off [fixed]
tx-gso-partial: off [fixed]
tx-tunnel-remcsum-segmentation: off [fixed]
tx-sctp-segmentation: off [fixed]
tx-esp-segmentation: off [fixed]
tx-udp-segmentation: off [fixed]
tx-gso-list: off [fixed]
tx-nocache-copy: off
loopback: off [fixed]
rx-fcs: off
rx-all: off
tx-vlan-stag-hw-insert: off [fixed]
rx-vlan-stag-hw-parse: off [fixed]
rx-vlan-stag-filter: off [fixed]
l2-fwd-offload: off [fixed]
hw-tc-offload: off [fixed]
esp-hw-offload: off [fixed]
esp-tx-csum-hw-offload: off [fixed]
rx-udp_tunnel-port-offload: off [fixed]
tls-hw-tx-offload: off [fixed]
tls-hw-rx-offload: off [fixed]
rx-gro-hw: off [fixed]
tls-hw-record: off [fixed]
rx-gro-list: off
macsec-hw-offload: off [fixed]
rx-udp-gro-forwarding: off
hsr-tag-ins-offload: off [fixed]
hsr-tag-rm-offload: off [fixed]
hsr-fwd-offload: off [fixed]
hsr-dup-offload: off [fixed]

Not sure if this will solve the problem, but what I did is create a script using the System Daemon to automatically adjust the NIC settings I stated above when the server reboots.

I will monitor for a week and if no further crashes, will update here