惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Security @ Cisco Blogs
The Last Watchdog
The Last Watchdog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
aimingoo的专栏
aimingoo的专栏
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
PCI Perspectives
PCI Perspectives
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
月光博客
月光博客
V
Visual Studio Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
Tailwind CSS Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
L
LangChain Blog
B
Blog RSS Feed
小众软件
小众软件
N
News | PayPal Newsroom
Attack and Defense Labs
Attack and Defense Labs
Microsoft Azure Blog
Microsoft Azure Blog
V
Vulnerabilities – Threatpost
The Hacker News
The Hacker News
T
Tor Project blog
A
Arctic Wolf
Jina AI
Jina AI
Hacker News: Ask HN
Hacker News: Ask HN
F
Fortinet All Blogs
Cloudbric
Cloudbric
S
Secure Thoughts
L
LINUX DO - 热门话题
博客园 - 司徒正美
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
S
Security Affairs
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
J
Java Code Geeks
P
Privacy International News Feed
AWS News Blog
AWS News Blog
S
Securelist
TaoSecurity Blog
TaoSecurity Blog
AI
AI
O
OpenAI News
C
Cyber Attacks, Cyber Crime and Cyber Security
K
Kaspersky official blog
T
The Blog of Author Tim Ferriss
大猫的无限游戏
大猫的无限游戏
Google DeepMind News
Google DeepMind News
Know Your Adversary
Know Your Adversary
P
Palo Alto Networks Blog
T
Tenable Blog
Last Week in AI
Last Week in AI
WordPress大学
WordPress大学
S
SegmentFault 最新的问题

Proxmox Support Forum

[SOLVED] - Github Auth for Mirrors-Kernel Repo? [Automation] Mass migration tool for MS Win11/Server Proxmox GUI hang - not response is it possible to reject or quarantine spam based on conditions I set ? The PVENode task list in PVE9 is partially obscured due to the terminal font being too large. About 100% error reporting due to pveproxy.service hooks Kubernetes overlay networking breaks when upgrading from PVE 9.1 to PVE 9.2.3 Zentraler Speicher No space left on device Combine datastore and direct file archival to tape Kernel panic VFS: Unable to mount root fs on unknown-block (0,0) sobald ein 7.x Kernel verwendet wird. How to migrate disk of a VM from one ZFS to another Windows Server 2025 fails to boot after PVE 9.2 / Linux 7.0 Kernel upgrade Cannot Install Proxmox on T610 Poweredge with H700 PERC card sdn Config. gateway not reachable How to safely change domain/FQDN? Welche Filterquote erreicht ihr? NFS Share status unknown on 2 of 5 nodes Can't connect to PVE9 consoles [solved] Can't connect to PVE9 consoles [solved] [SOLVED] - Use secondary network for PVE commands Created cluster, one node storage gone BUG: proxmox mail gateway FROM = null bypass spam filtering Moving existing PBS from VMWare workstation to PVE cluster Does eBGP SDN fabric support external peering? Bug: PDM 1.1 not recognizing valid license status Proxmox GUI hang - not response PVE crashes unexpectedly Proxmox Backup Server 4.2 released! Advice [META] Links on Proxmox Forum Website Hardwarer oder Software RAID Joining a cluster with already created guests VM PDM missing backup jobs from PVE / Log retention Remove VM.Monitor from all users/roles, PVE 9.2 Proxmox Freezing (new instalation) 9.2.2 - Intel 12700T No Web gui and random connection reset by peer [SOLVED] - i40e module for X710 Intel NIC Dutch Proxmox Day 2026 How pools use the space Corosync initiiert Reboot trotz Verfügbarkeit der Systeme Opt-in Linux 7.0 Kernel for Proxmox VE 9 available After PVE 8to9 upgrade, unable to check guest fs freeze status Problem with MegaRAID SAS3508 controller proxmox-kernel-7.0.2-6-pve failing network service Auto sync guest time after rollback of VM snapshot with RAM/state Broadcom BCM57504 (100G) bnxt_en TX timeout and NIC reset on Proxmox 8.1.5 — while BCM57414 (25G) works fine on same host QEMU 11.0 available on pve-test and pve-no-subscription as of now 350 MPM Solventless Lamination Machine for High-Speed Flexible Packaging Making sense of NVMe zfs and SMART errors [SOLVED] - PVE loses network connection after kernel upgrade to proxmox-kernel-7.0.0-3-pve [SOLVED] - Remove or reset cluster configuration. Proxmox 8.4.1 Fresh Install BCM57416 10G Ethernet Adapter Not Recognized PDM 1.1.1 unable to add AD realm with anonymous search [TUTORIAL] - Developer Workstation (Proxmox-VE 9) with cinnamon (LMDE7) SDN zone shows "pending" on peer nodes after node reboot (9.2.x) Cluster not quorate - extending auth key lifetime! Proxmox not rebooting properly (SOLVED) Proxmox 9 Stuck on loading initial ramdisk With new HA-Disarm Feature is there a Documentation for NUT Setup on Clusters? Proxmox 8.3 Installation Issue on ProLiant DL380 Gen9 Cluster networking setup LXC System images unavailable [SOLVED] - Fix: NVIDIA Drivers Failing after upgrade to Proxmox 9.2.2 (Kernel 7.0.2-6-pve) / NovaCore Conflict Install NUT directly on Proxmox VE and control guests from here driver usb for windows 7 System startup error and no network: Failed to start ifupdown2-pre.service - Helper to synchronize boot up for ifupdown. PBS backup space grow up constantly Proxmox Datacenter Manager 1.1 released! IPv4 not available in newly created VM Recommended Setup for Offsite Proxmox Backups? Hetzner Storage Box & Remote PBS Challenges duplicate, please delete this passthrought an USB device "by ID" to CT PDM Installer Freezes at 66% Tried PDM for the first time (version 1.1) - had issues PDM 1.1 automated install Suche Server-Provider für Proxmox connecting sdn to edge firewall SDN, IPAM & DHCP Migrating from read-only file system Ubuntu 26.04 installation fails for unknown reason Status Unbekannt nach Cluster Join Installing Proxmox Backup Server on Mac Mini (Late 2012) kernel 7.0 performance issue with zfs pools PVE becomes unreachable via ethernet but OS is running [SOLVED] - New 9.2 install - can't find 7.0.2-6-pve , not all the time [SOLVED] - Backup and dedupe a VM with LUKS Gibt es mit PVE 2.x ggf. Änderungen bei der RAM-Nutzung, bzw. deren Anzeige bei VMs? I need help for setting up backup solution Way more NAGware, very little functionality, bugs galore Root squashing virtiofsd with --uid-map Intel ixgbe Driver Update Fail Help to fix Proxmox access issues after power cut Passkey Login (not 2FA) Roblox VM detection - can be overcome? [TUTORIAL] - ZFS-Autosnaptshot inkl. Rollback und Daten direkt recovern (Windows/Linux) How to stop PVE Kernel upgrade [SOLVED] - very long waiting to log in to lxc debian 11 ssh [TUTORIAL] - Configuring Fusion-Io (SanDisk) ioDrive, ioDrive2, ioScale and ioScale2 cards with Proxmox Increase maximum USB devices in vm.conf
ceph-osd crashes with kernel 6.17.2-1-pve on Dell system
invalid@exam · 2026-05-29 · via Proxmox Support Forum

Hey! Recently i upgraded one of the three running nodes in a cluster to 6.17.2-1-pve kernel, ceph version remains the same on all hosts (19.2.3).

When server rebooted i noticed that instantly ceph-osd processes were crashing:

Code:

ceph-osd[10805]: ./src/common/HeartbeatMap.cc: 85: ceph_abort_msg("hit suicide timeout")

And kernel did throw these stack traces:

Code:

kernel: sd 0:2:0:0: [sda] tag#616 page boundary ptr_sgl: 0x00000000df48bcb9
kernel: BUG: unable to handle page fault for address: ff685a6f8dd63000
kernel: #PF: supervisor write access in kernel mode
kernel: #PF: error_code(0x0002) - not-present page
kernel: PGD 100000067 P4D 100874067 PUD 100875067 PMD 108abd067 PTE 0
kernel: Oops: Oops: 0002 [#1] SMP NOPTI
kernel: CPU: 81 UID: 0 PID: 1012 Comm: kworker/81:1H Tainted: P S         OE       6.17.2-1-pve #1 PREEMPT(voluntary)
kernel: Tainted: [P]=PROPRIETARY_MODULE, [S]=CPU_OUT_OF_SPEC, [O]=OOT_MODULE, [E]=UNSIGNED_MODULE
kernel: Hardware name: Dell Inc. PowerEdge R660xs/00NDRY, BIOS 2.7.5 07/31/2025
kernel: Workqueue: kblockd blk_mq_run_work_fn
kernel: RIP: 0010:megasas_build_and_issue_cmd_fusion+0xeaa/0x1870 [megaraid_sas]
kernel: Code: 20 48 89 d1 48 83 e1 fc 83 e2 01 48 0f 45 d9 4c 8b 73 10 44 8b 6b 18 4c 89 f9 4c 8d 79 08 45 85 fa 0f 84 fd 03 00 00 45 29 cc <4c> 89 31 48 83

kernel: RSP: 0018:ff685a6fa0b0fb50 EFLAGS: 00010206
kernel: RAX: 00000000fe298000 RBX: ff42339b0e6b2cc0 RCX: ff685a6f8dd63000
kernel: RDX: ff685a6f8dd63008 RSI: ff42339b0e6b2b88 RDI: 0000000000000000
kernel: RBP: ff685a6fa0b0fc20 R08: 0000000000000200 R09: 0000000000001000
kernel: R10: 0000000000000fff R11: 0000000000001000 R12: 0000000000101000
kernel: R13: 0000000000102000 R14: 0000000009a00000 R15: ff685a6f8dd63008
kernel: FS:  0000000000000000(0000) GS:ff4233da0e986000(0000) knlGS:0000000000000000
kernel: CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
kernel: CR2: ff685a6f8dd63000 CR3: 0000000133351004 CR4: 0000000000f73ef0
kernel: PKRU: 55555554
kernel: Call Trace:
kernel:  <TASK>
kernel:  megasas_queue_command+0x122/0x1d0 [megaraid_sas]
kernel:  scsi_queue_rq+0x409/0xcc0
kernel:  blk_mq_dispatch_rq_list+0x121/0x740
kernel:  ? sbitmap_get+0x73/0x180
kernel:  __blk_mq_sched_dispatch_requests+0x408/0x600
kernel:  blk_mq_sched_dispatch_requests+0x2d/0x80
kernel:  blk_mq_run_work_fn+0x72/0x90
kernel:  process_one_work+0x188/0x370
kernel:  worker_thread+0x33a/0x480
kernel:  ? __pfx_worker_thread+0x10/0x10
kernel:  kthread+0x108/0x220
kernel:  ? __pfx_kthread+0x10/0x10
kernel:  ret_from_fork+0x205/0x240
kernel:  ? __pfx_kthread+0x10/0x10
kernel:  ret_from_fork_asm+0x1a/0x30
kernel:  </TASK>

(Stack traces can be seen for all 5 ceph-osd disks i have in a system, the stack trace is the same only the drive letter changes)

I tried restarting the host, recreating the osds - nothing helped, only thing that helps is to boot into older installed kernel i have on the system - 6.14.11-3-pve, when i boot into that kernel - everything works like a charm.

I did ran the memtest on the system just to make sure its not some hardware issue, aswell as reseated the cables for backplane and stuff.
Here is some info about the hardware:
Dell Inc. PowerEdge R660xs
Bios: 2.7.5 (newest)
Raid controller: PERC H755N (version: 52.30.0-6115 - newest) - Disks are passed through to the system as NON-RAID disks.
ceph version 19.2.3 (2f03f1cd83e5d40cdf1393cb64a662a8e8bb07c6) squid (stable)
pve-manager/9.0.18/5cacb35d7ee87217 (running kernel: 6.14.11-3-pve)
While reading forums i noticed some threads regarding people having issues with newer kernel on Dell systems (maybe related?)

Hello,
we have the syame problem on our HPE Hosts.
Booting on older Kernel 6.14 resolves the Problem.
i have currently one node up with new kernel (6.17) for debugging purposes if someone need outputs or logs...

Is there anyone who can help with the problem? Or is downgrading the kernel back to 6.14 the solution?

For now i've only found downgrading the kernel as the solution, but the newer kernels have to include some sort of fix, otherwise we all expriencing the issue will be stuck with the older kernel.

For now i've only found downgrading the kernel as the solution, but the newer kernels have to include some sort of fix, otherwise we all expriencing the issue will be stuck with the older kernel.

We're currently testing a new kernel with a larger set of changes. Nothing specific for megaraid_sas - but quite some changes in the scsi subsystem.
It's currently available in the pbs-test repository (and will soon be available for pve-test as well).

A quick search online did not show too many hits for this particular stacktrace - only something remotely related for a much older kernel on SLES:
https://stgsupport.stgscc.suse.com/...ontrollers-randomly-crash-on-boot?language=de

Sadly we could not yet reproduce the issue and don't have a matching system.

If you can trigger the issue reliably (in a non-critical environment) - trying the new kernel when it's available and/or setting
`smp_affinity_enable=0` for the module might help in getting this narrowed down and fixed.

Thanks for the report in any case!

A similar trace was also reported in the general kernel 6.17 announcement thread:

sd 0:2:1:0: [sda] tag#4057 page boundary ptr_sgl: 0x00000000ba1fad69[ 28.571202] BUG: unable to handle page fault for address: ff72bd070403c000[ 28.571210] #PF: supervisor write access in kernel mode[ 28.571216] #PF: error_code(0x0002) - not-present page[ 28.571222] PGD 100000067 P4D 100304067 PUD 100305067 PMD 12ddba067 PTE 0[ 28.571232] Oops: Oops: 0002 [#1] SMP NOPTI[ 28.571240] CPU: 5 UID: 0 PID: 1205 Comm: kworker/u128:4 Tainted: P O 6.17.2-2-pve #1 PREEMPT(voluntary) [ 28.571250] Tainted: [P]=PROPRIETARY_MODULE, [O]=OOT_MODULE[ 28.571256] Hardware name...

i just updated kernel to 6.17.2-2 (is this the kernel you mentioned?). But same behaviour as with 6.17.2-1.
When i set the noin flag on the ceph cluster and reboot the host, the OSDs are show as UP/OUT. As soon as i set it to in it goes down...
in the ceph-osd log i see some entires for "transitioning to primary" then "transitioning to stray" and then i get spammed with
7e377f26b6c0 1 heartbeat_map is_healthy 'OSD: osd_op_tp thread 0x7e3762cb56c0' had timed out after 15.000000954s
until i set it to out again. After that i cant restart the service an have to reboot the server.

After booting to 6.14.11-4 i can set the osds to in without problems...

Last edited:

i just updated kernel to 6.17.2-2 (is this the kernel you mentioned?).

no that would be proxmox-kernel-6.17.4-1-pve - I'll post here when it's available in the public pve repos as well (currently only on pbs-test)

But thanks for the test - at least it rules out that the regression came in between 6.17.2-1 and 6.17.2-2

When i set the noin flag on the ceph cluster and reboot the host, the OSDs are show as UP/OUT. As soon as i set it to in it goes down...
in the ceph-osd log i see some entires for "transitioning to primary" then "transitioning to stray" and then i get spammed with
7e377f26b6c0 1 heartbeat_map is_healthy 'OSD: osd_op_tp thread 0x7e3762cb56c0' had timed out after 15.000000954s

I don't think it's a ceph-specific problem - the other reporter in the general thread ran into the kernel trace by running `proxmox-boot-tool refresh` (which doesn't do much I/O either)

Same here, but with fully updated PBS 4.1 and enterprise repos. The server is an HP DL360 Gen10+, and the RAID controller an HP MR416i-a Gen10+.

When the server hits the issue, we have to do a hard reset since some processess get stuck in D state and it's impossible to finish a graceful reboot.

Currently rolling back to 6.14 and see how it works.

6.14 resolved it for us too. We're doing backups since Monday without issues.

We had this issue also. Rebooted to 6.14 and removed 6.17

For those who are looking for the commands:

Code:

# Boot to 6.14

# Rollback these 2 packages to remove dependency on 6.17 kernel
sudo apt install proxmox-default-kernel=2.0.0 proxmox-ve=9.0.0

sudo apt purge proxmox-kernel-6.17 proxmox-kernel-6.17.4-2-pve-signed

Is there an issue opened for this ? I would like to have an official thread to follow up

i get the kernel 6.17.4-2 listed:
1769705748167.png
anyone tested yet?

no that would be proxmox-kernel-6.17.4-1-pve - I'll post here when it's available in the public pve repos as well (currently only on pbs-test)

But thanks for the test - at least it rules out that the regression came in between 6.17.2-1 and 6.17.2-2

I don't think it's a ceph-specific problem - the other reporter in the general thread ran into the kernel trace by running `proxmox-boot-tool refresh` (which doesn't do much I/O either)

"I'm having the same issue on a Hp dl380 Gen11 Raid mr416i-o, NVM Micron 7400, Rolling back to 6.14 fixed it for me as well."

Same here: ProLiant DL380 Gen10 Plus, HPE MR416i-p Gen10+

Fresh install, when adding OSDs, some show up as "bluestore" and seem to be ok, some show up as "filestore" - very confusing until I found this thread.

Downgrading to 9.0 solved the issue.

Would be nice to have a more "prominent" announcement somewhere, since this is a very frustating experience (aka "showstopper") for new users. Or, maybe have at least some more feedback on progress regarding this issue.

hd--

Proxmox Staff Member

There is the new Kernel version 6.17.9-1 in the no-subscription repository, which might be worth a test

I hit this issue with a new build on HPE DL360 8SFF GEN11 servers on the latest kernel with all available updates in the no-subscription repository installed. Rolling back to 6.14 using the instructions in #11 resolved it.

Has anyone performed testing on kernel version 6.17.13?

Tested 6.17.4-2, not working.

Is there a safe way to switch from subscription to no-subscription and back? I would be willing to test, but would like to go back to subscription afterwards .

Is there a safe way to switch from subscription to no-subscription and back? I would be willing to test, but would like to go back to subscription afterwards .

You can switch to no-subscription repository, do apt update, install the desired kernel doing apt install proxmox-kernel-6.xx (input your kernel version here), then go back to enterprise repository and do apt update again. Make sure you only install the desired kernel and anything else when in no-subscription.

If the new kernel doesn't work for you, you can either pin an older kernel or directly remove it using apt remove proxmox-kernel-6.xx