Skip to content
DnsLister Forum

Where domain hunters compare notes

Raven 2 (Picasso) amdgpu/Display Core lockup under high thermal load — temperature is the trigger, upstream GFXOFF workaround does not resolve it — HP 245 G7, Ryzen 3 3250U

System Information

  • Laptop: HP 245 G7 Notebook PC
  • CPU/APU: AMD Ryzen 3 3250U with Radeon Vega 3 Graphics (Raven 2 / Picasso, PCI ID 1002:15d8, rev c4)
  • RAM: 16 GB DDR4 (dual channel)
  • Storage: NVMe SSD
  • BIOS: F.77 (latest available, dated 2025-11-28)
  • Battery: Replaced (Neotek HT03XL, 41.04 Wh, calibrated, health OK)
  • Distro: CachyOS (Arch-based)
  • Kernel: 7.2.6-1-cachyos (also tested 6.18.52-1-cachyos-lts)
  • Mesa: 26.2.3-arch3.1
  • Desktop: X11 (LXQt with Openbox, no compositor)
  • GPU driver: amdgpu (kernel)
  • Display: eDP-1, ChiMei InnoLux 0x14c3, 1366×768 @ 75 Hz

Problem Description

The system completely freezes (screen locked, audio loop, hard reset required) when playing GPU-intensive games via Steam Proton. Reproduced with:

  • Detroit: Become Human (Decima Engine)
  • Resident Evil 4 Remake (RE Engine)

The freeze is not a thermal shutdown. It is an amdgpu/Display Core (DC) lockup. The kernel stack trace captured at the moment of freeze:

INFO: task kworker/u16:0:12 blocked for more than 122 seconds.

Workqueue: events_unbound commit_work

Call Trace:

__schedule

schedule

schedule_timeout

dma_fence_default_wait

dma_fence_wait_timeout

drm_atomic_helper_wait_for_fences

commit_tail

process_scheduled_works

worker_thread

kthread

amdgpu … dm_read_reg_func

amdgpu … generic_reg_get

amdgpu … dc_stream_get_scanoutpos

Also captured:

amdgpu 0000:04:00.0: Dumping IP State

pipewire[827]: spa.alsa: front:1p: (0 suppressed) snd_pcm_avail after recover: Tubería rota

The Display Core is stuck waiting on a fence during an atomic commit. Audio fails as a consequence (shared resources with the display).

Key Finding: Temperature is the Trigger

Extensive testing has shown that the lockup is strongly temperature-dependent.

Configuration Temp Duration Result
GFXOFF ON (default), ryzenadj aggressive limits (6W, 65 °C) 65-76 °C 35+ min ✅ Stable
GFXOFF ON (default), moderate limits ~87 °C ~15 min ❌ Freeze
GFXOFF ON (default), no limits 90-96 °C ~5-22 min ❌ Freeze
GFXOFF OFF (0xfff73fff), no temp limit 100-103 °C N/A ❌ Thermal shutdown

Conclusion: With GFXOFF enabled and identical Display Core configuration, the system is stable at low temperatures and consistently locks up above ~90 °C. The GFXOFF/DC bug is real and confirmed by the stack trace, but temperature is the trigger that exposes it.

Important Context

  1. The upstream workaround b35eb912 is present in the tested kernel. Confirmed via strings on amdgpu.ko:
  2. textgfx_v9_0_ring_begin_use_compute gfx_v9_0_ring_end_use_compute This is the patch drm/amdgpu/gfx9: manually control gfxoff for CS on RV (integrated in Linux 6.12.17). It protects compute queues during GFXOFF transitions. The lockup persists despite this patch being present.

  3. Overdrive is not implemented on Raven 2. Kernel logs show: textamdgpu: pp_odn_edit_dpm_table was not implemented. amdgpu: pp_dpm_get_sclk_od was not implemented. amdgpu: pp_od_clk_voltage is not accessible if power_dpm_force_performance_level is not in manual mode! Attempts to enable Overdrive via ppfeaturemask=0xffffffff or LACT trigger immediate instability.

  4. HP firmware blocks the fan when no battery is detected. With battery installed and calibrated, the fan works. HP firmware-related ACPI errors: textACPI BIOS Error (bug): Could not resolve symbol [\_SB.PCI0.GPP2.BCM5], AE_NOT_FOUND ACPI Error: AE_NOT_FOUND, During name lookup/catalog ACPI: thermal: [Firmware Bug]: No valid trip points! hp-wmi hp-wmi: Failed to apply initial fan settings: -22 ACPI Error: AE_AML_BUFFER_LIMIT hp_bioscfg: Returned error 0x3, "Invalid command value/Feature not supported"

Configurations Tested

ppfeaturemask GFXOFF Stutter DCS OD Result
0xfff7bfff (kernel default) ON ON OFF OFF ❌ Freezes
0xfffd3fff OFF ON OFF OFF ✅ Stable, 100 °C, thermal shutdown
0xffff7fff OFF ON ON ON ✅ Stable, 100+ °C
0xfff5bfff ON OFF OFF OFF ❌ Freezes
0xfff73fff OFF ON OFF OFF ✅ Stable, 100+ °C, thermal shutdown
0xffffffff (Overdrive ON) ON ON ON ON ❌ Immediate freeze at menu
Default + amdgpu.dcdebugmask=0x8 (DC clock gating off) ON ON OFF OFF ❌ Freezes
Default + amdgpu.dcdebugmask=0x10 (PSR off) ON ON OFF OFF ❌ 22 min then freeze at 96 °C
Default + dcdebugmask=0x10 + 60 Hz ON ON OFF OFF ❌ Freezes
Default + dcdebugmask=0x40 (MPO off) ON ON OFF OFF Not conclusively tested
GFXOFF ON + ryzenadj savings (6W, 65 °C) ON ON OFF OFF ✅ 35 min stable

Note: ppfeaturemask=0xfff7bfff is the kernel default for Raven 2 on kernel 7.x. It enables GFXOFF and disables Overdrive and DCS.

Other Parameters Tested (No Effect)

  • amdgpu.sg_display=0 → no effect
  • amdgpu.dpm=0 → prevents boot
  • amdgpu.gpu_recovery=1 → recovery does not complete before freeze
  • 60 Hz instead of 75 Hz → still freezes
  • Early KMS (amdgpu in initramfs MODULES) → no effect
  • Reinstalled CachyOS from scratch → no effect
  • Tested linux-cachyos-lts 6.18.52 → same behavior

Hardware Maintenance Performed

  • Laptop opened and inspected.
  • Heat pipe found slightly bent near the CPU contact point. Gently straightened to reduce wobble.
  • Heatsink cleaned (dust on fins).
  • Thermal paste replaced (LK-17, nominal 17 W/mK, applied ~3 months ago, previously degraded by multiple 100+ °C sessions).
  • New battery installed (HT03XL, calibrated, health 100%).
  • Idle temperatures after maintenance: ~40 °C (Tctl), ~40 °C (amdgpu edge). Previously 46-62 °C.

Sensor Data

Sensors available on this system:

  • k10temp (hwmon4): Tctl — no offset on Raven 2 mobile (Tctl ≈ Tdie ≈ amdgpu edge, verified in idle)
  • amdgpu (hwmon6): edge
  • acpitz (hwmon1): temp1
  • nvme (hwmon3): Composite
  • hp (hwmon5): no sensor exposed
  • BAT0 (hwmon2): battery temp

Important: On this APU, Tctl does not have an artificial offset. Values reported are real die temperature.

What We Believe

The GFXOFF/Display Core bug is real and confirmed by the stack trace. The upstream compute-queue workaround (b35eb912) is present but does not prevent the lockup.

Our hypothesis: At high temperatures (≥90 °C), the SMU/firmware may change power states, clock domains, or voltage thresholds in a way that exposes a synchronization failure in the Display Core's atomic commit path. This may be a second, temperature-dependent bug distinct from the compute-queue issue that b35eb912 addresses.

We are not claiming this as fact. We are looking for confirmation or correction from developers.

Reproduction Steps

  1. Boot with default ppfeaturemask=0xfff7ffff (GFXOFF enabled, default for Raven 2).
  2. Ensure system is at room temperature (~40 °C idle).
  3. Launch Detroit: Become Human or RE4 Remake via Steam Proton (Proton Experimental or GE-Proton tested).
  4. Play at 720p/Low settings.
  5. Observe:
    • At 65-76 °C: system remains stable for 35+ minutes.
    • Above 90 °C: system freezes with the stack trace above.

Attached Logs

dmesg (freeze event):

amdgpu 0000:04:00.0: Dumping IP State

INFO: task kworker/u16:0:12 blocked for more than 122 seconds.

Workqueue: events_unbound commit_work

Call Trace:

__schedule

schedule

schedule_timeout

dma_fence_default_wait

dma_fence_wait_timeout

drm_atomic_helper_wait_for_fences

commit_tail

process_scheduled_works

worker_thread

kthread

amdgpu … dm_read_reg_func

amdgpu … generic_reg_get

amdgpu … dc_stream_get_scanoutpos

amdgpu module parameters (current, working config):

ppfeaturemask = 0xfff7ffff dcdebugmask = 0x10 (PSR off) 

Module strings confirming upstream patch presence:

gfx_v9_0_ring_begin_use_compute
gfx_v9_0_ring_end_use_compute

Request for Help

We are asking the AMD DRM developers:

  1. Is this a known second issue with Raven 2/Picasso, distinct from b35eb912?
  2. Is there a way to make the Display Core more conservative (beyond dcdebugmask=0x10 and 0x8) to prevent the lockup without disabling GFXOFF?
  3. Would a patch that controls GFXOFF more granularly during the DC atomic commit path solve this?
  4. Is there a test case or diagnostic we can run to isolate the exact failure mechanism?
  5. Would disabling specific DC features (e.g., dcdebugmask=0x40 for MPO) prevent this?

We are willing to test patches, provide further diagnostics, or run custom kernels.

Environment Details

  • CachyOS base with CachyOS kernel 7.2.6-1
  • LXQt desktop, X11, no compositor
  • All packages up to date as of 2026-09-24
  • Reproducible across two kernels (7.2.6-1-cachyos and 6.18.52-1-cachyos-lts)

Final Note

This issue has been ongoing for several weeks. We have tried every reasonable software workaround. The only stable configurations are:

  • GFXOFF ON + temperature < 80 °C (works, but requires aggressive power limits and low FPS)
  • GFXOFF OFF (stable but causes thermal shutdown at 100+ °C)

We are looking for a real fix, not a workaround.

https://pastebin.com/tfpHCEa6

Source: r/linux_gaming · by /u/Economy-Ad3574

Leave a Reply

Your email address will not be published. Required fields are marked *