Windows Server - HP Proliant ML350 server running SBS 2003 FREEZES with Event ID 7011

Asked By Abid Parwaz on 25-Jun-10 05:46 PM
Hi,

We have a SBS 2003 server that has some serious issues for few weeks now. The
problem is that server freezes for a short span of time and comes back to
life. Freeze happens many times every hour and lasts 50-60 secs each time and
then come back to life. During the freeze:

• Users get the popup balloon in Outlook telling them that Exchange is trying
to retrieve data.
• Shared drives cannot be opened.
• Softwares retrieving data from the server also appear hang/freeze.
• Remote Desktop connection to the server continues but cannot perform any
actions (only mouse pointer moves but cannot select anything).
• During the time of freeze following Events are generated in System log:
 Going through the logs Event ID: 7011 always seems to be recorded around
the time the freeze happens. It says, 'Timeout (60000 milliseconds) waiting
for a transaction response from the NtFrs service'.
Also in Applications Log This error is found: Event Source: MSSQLSERVER
The Scheduler 0 appears to be hung. SPID 0, ECID 0, UMS Context 0x03640BC8

• When the server is booted it takes over 40 mins on 'Applying computer
settings'.
• When the server is restarted or shutdown it takes about 15-20 mins.

A bit about the server:

OS: Windows SBS 2003 Premium Edition – SP2, with Exchange 2003 SP2, SQL 2000,
SharePoint Services 2.0
Hardware: HP ProLiant ML350 G4
Smart Array 641 Controller
6 x 146.8GB HDD (10K rpm) configured into 3 Logical drives
RAID 1+0 is used
x1 Intel Xeon 3.0GHz Processor
4GB of RAM installed – SBS 32-bit only uses 3.5GB

Hardware diagnostics (using HP diag tools) show no issues. Green LEDs are
shown on all 6 Hard Drives. I have already updated all device drivers and
system firmware using HP tools. HP Array Configuration Utility shows no
issues at all.

Please help!
Ramu Soft replied to Abid Parwaz on 26-Jun-10 10:31 AM

In my experience I have seen several server hang issues. In most cases troubleshooting ended at HDDs. Though the status of HDDs shows green there might be some problem with them and in turn delaying the write/read operations.


Do one thing. Go to your perfmon and verify the "average disk queue length" counter values. If they are above 10, then you have something to worry about your disks.


Hope this information helps you.


Thanks,

Abid Parwaz replied to Ramu Soft on 28-Jun-10 05:46 AM

Thanks,


I have updated the StorPort driver according to Microsoft KB932755 at http://support.microsoft.com/kb/932755 HP Advisory also had suggested installing this update on ProLiant ML servers.


The Fixes list also includes 'Server becomes unresponsive or blue screen errors upon shutdown etc. Installation went well and the reboot took nearly 2 hrs to complete shutdown and reboot. Its too soon to say if the problem has been resolved. I am constantly monitoring the logs. I hope and strongly wish that the problem has been resolved.


The Performance Monitor was already enabled and it shows a graph of 3 things, Page File, Avg. Disk Queue Length and Processor. Now that the server has been restarted and there is no heavy load so all seems fine. At the moment Avg. Disk Queue Length reaches an Avg value of 1.8.


I'll continue monitoring and update this thread soon.


Many thanks!!!!

Abid Parwaz replied to Abid Parwaz on 01-Jul-10 05:13 PM

Hi,


I am afraid the problem is back ... It never went away in the first place... the server was rebooted so all services were refreshed... This makes me believe that all these errors are somehow related to disk I/O ( or may be) memory operations and problem must be with the hardware. Problem could be with the Array Controller 641 trying to read data from the disks. Or could be with any of the 6 hard drives. But all comes healthy :( during HP diagnostic tools. System Idle Process in Task Manager shows cpu usage of > 90% but in performance tab it shows the very normal cpu usage.

I can paste here the HP Diagnostic results which I generated using HP SmartStart CD by booting system from it.

Don't mean to increase the length of the thread but really need to find a solution.


Thanks for taking out your valuable time!!

Abid Parwaz replied to Abid Parwaz on 01-Jul-10 05:17 PM

I have already tried it but no happiness :( ....


  • Autodisconnect value is already set to 0xfffff, mentioned  (they said its th highest value allowed for a DC)
  • ServicesPipeTimeout value is already set to 60000
  • This server is the only DC on the domain and I had turned off NtFrs service but no use. NtFrs is only a symptom of some other hardware/software disease. Turning it off only turns off Event ID 7011 but the freeze doesn't go away.

The problem has got worse with installing SortPort hotfix KB932755. Now the Yosemite backup software does not finish the over-night backup. After backing up only 41% out of 275Gb of data it stops with an error. The error reads, 'The status of the SCSI bus is invalid. Please check the device and connections''. I have already tested the Tape drive all is fine. No errors, nothing.


I am thinking of uninstalling the SortPort driver this saturday and reboot the server. and keep looking for the fix.

btw, I have configured Performance Monitor to log Memory, Processor, LogicalDisk, PhysicalDisk counters but the problem is there is no timestamp on it. From the Event Viewer I know the exact time of the freeze happening but I cant compare it with the Perfmon log.


I am also going to tuning up the SQL instances to use a set amount of RAM.


There  were users here on this forum with similar issue on ProLiant ML350. Can anyone please point out what was their finding.... Plz help!!!

Abid Parwaz replied to Abid Parwaz on 01-Jul-10 05:19 PM
I have been observing the Performance Monitor tool. Memory/%Pages/sec  and PhysicalDisk Avg Disk Queue length counters are constantly high. Here is the screenshot!

http://img819.imageshack.us/img819/7651/perfmon.jpg 

Also the Memory\%Available Megabytes is constantly hitting the full scale. 4GB memory is installed but System > properties only show 3.50GB. In Taskbar PF usage was 3.47GB. I killed java and spiceworks process and it came down to 3.27GB but the Memory/%Available MB counter is still the same.


May be the system is RAM starved. There is an article on Microsoft website that talks about the /PAE switch. But I have further read that its enabled in SBS 2003 by default and should not be added to Boot.ini. I have never seen task Manager using any memory over 3.5 GB. May be the rest of teh memory is used by other device/mappings etc. OR its possible that my server is actually using only 3.5GB and wasting 500MB. Shall I enable the /PAE switch on boot.ini??


Here is an HP Customer Advisory http://h20000.www2.hp.com/bizsupport/TechSupport/Document.jsp?lang=en&cc=us&objectID=c00883105&jumpid=reg_R1002_USEN titled as {Advisory: (Revision) Certain Operating Systems May Not Report All Installed Memory for HP ProLiant Servers with 4 GB or Greater of Memory if PAE Mode Is Not Enabled} which is true to my situation.