Creation Zone

  • Subscribe to our RSS feed.
  • Twitter
  • StumbleUpon
  • Reddit
  • Facebook
  • Digg

Wednesday, 9 April 2008

Running Batch Workloads on Sun's CMT Servers

Posted on 09:24 by Unknown
(Originally posted on blogs.sun.com at:
http://blogs.sun.com/mandalika/entry/running_batch_workloads_on_sun
)

Ever since Sun introduced Chip Multi-Threading (CMT) hardware in the form of UltraSPARC T1's T1000/T2000, our internal mail aliases were inundated with variety of customer stories, majority of those go like 'batch jobs are taking 12+ hours on T2000, where as it takes only 3 or 4 hours on US-IV+ based v490'. Even after two and half years since the introduction of the revolutionary CMT hardware, it appears that majority of Sun customers are still under the impression that Sun's CMT systems like T2000, T5220 are not capable of handling CPU intensive batch workloads. It is not a valid concern. CMT processors like UltraSPARC T1, T2, T2 Plus can handle batch workloads just as well like any other traditional/conventional processor viz. UltraSPARC-IV+, SPARC64-VI, AMD Opteron, Intel Xeon, IBM POWER6. However CMT awareness and little effort are required at the customer end to achieve good throughput on CMT systems.

First of all, the end users must realize the fact that the maximum clock speed of the existing CMT processor line-up (UltraSPARC T1, UltraSPARC T2, UltraSPARC T2 Plus) is only 1.4 GHz; and on top of that each strand (individual hardware thread) within a core shares the CPU cycles with the other strands that operate on the same core (Note: each core operates at the speed of the processor). Based on these facts, it is no surprise to see batch jobs taking longer times to complete when only one or a very few single-threaded batch jobs are submitted to the system. In such cases, the system resources are fairly under-utilized in addition to the longer elapsed times. One possible trick to achieve the required throughput in the expected time frame is to split up the workload into multiple jobs. For example, if an EDU customer needs to generate 1000 transcripts, the customer should consider submitting 4 individual jobs with 250 transcripts each or 8 jobs with 125 transcripts each rather than submitting one job for all 1000 transcripts. Ideally the customer should observe the resource utilization (CPU%, for example); and experiment with the number of jobs to be submitted until the system achieves the desired throughput within the expected time frame.

Case study: Oracle E-Business Suite Payroll 11i workload on Sun SPARC Enterprise T5220

In order to prove that the aforementioned methodology works beyond a reasonable doubt, let's take Oracle's E-Business Suite 11.5.10 Payroll workload as an example. On a single T5220 with one 1.4 GHz UltraSPARC T2 processor, acting as the batch, application and database server, 4 payroll threads generated 5,000 paychecks in 31.53 minutes of time consuming only 6.04% CPU on average. ~9,500 paychecks is the projected hourly throughput. This is a classic example of what majority of Sun's CMT customers are experiencing as of today i.e., longer batch processing times with little resource consumption. Keep in mind that each UltraSPARC T2 and UltraSPARC T2 Plus processors can execute up to 64 jobs in parallel (on a side note, UltraSPARC T1 processor can execute up to 32 jobs in parallel). So to put the idling resources for effective use, there by to improve the elapsed times and the overall throughput, few experiments were conducted with 64 payroll threads and the results are very impressive. With a maximum of 64 payroll threads, it took only 4.63 minutes to process 5,000 paychecks at an average of 40.77% CPU utilization. In other words, similarly configured T5220 can process ~64,700 paychecks at less than half of the available CPU cycles. Here is a word of caution: just because the processor can execute 64 threads in parallel, it doesn't mean it is always optimal to submit 64 parallel jobs on systems like T5220. Very high number of batch jobs (payroll threads in this particular scenario) might be an overkill for simple tasks like NACHA in Payroll process.

The following white paper has more detailed information about the nature of the workload and the results from the experiments with various number of threads for different components of the Oracle Applications' Payroll batch workload. Refer to the same white paper for the exact tuning information as well.

Link to the white paper:
     E-Business Suite Payroll 11i (11.5.10) using Oracle 10g on a Sun SPARC Enterprise T5220

Here is the summary of the results that were extracted from the white paper:

Hardware configuration

          1x Sun SPARC Enterprise T5220 for running the application, batch and the database servers
              Specifications: 1x 1.4 GHz 8-core UltraSPARC T2 processor with 64 GB memory

Software configuration

          Oracle E-Business Suite 11.5.10
          Oracle 10g R1 10.1.0.4 RDBMS
          Solaris 10 8/07

Results

Oracle E-Business Suite 11i Payroll - Number of employees: 5,000
Component#ThreadsTime (min)Avg. CPU%Hourly Throughput
Payroll process641.8790.56160,714
PrePayments640.2046.331,500,000
Ext. Proc. Archive641.9090.77157,895
NACHA80.052.526,000,000
Check Writer240.389782,609
Costing480.2332.51,285,714
Total or AverageNA4.63 min40.77%64,748


It is evident from the average CPU% that the Payroll process and the External Process Archive components are extremely CPU intensive; and hence take longer time to complete. That's the reason 64 threads were configured for those components to run at the full potential of the system. Light-weight components like NACHA need fewer threads to complete the job efficiently. Configuring 64 threads for NACHA will have a negative impact on the throughput. In other words, we would be wasting CPU cycles for no apparent improvement.

It is the responsibility of the customers to tune the application and the workload appropriately. One size doesn't fit all.

The Payroll 11i results on the T5220 demonstrate clearly that Sun's CMT systems are capable of handling batch workloads well. It would be interesting to see how well they perform against other systems equipped with traditional processors with higher clock speeds. For this comparison, we could use couple of results that were published by UNISYS and IBM with the same workload. The following table summarizes the results from the following two white papers. For the sake of completeness, Sun's CMT results were included as well.

Source URLs:
  1. E-Business Suite Payroll 11i (11.5.10) using Oracle 10g on a UNISYS ES7000/one Enterprise Server
  2. E-Business Suite Payroll 11i (11.5.10) using Oracle 10g for Novell SUSE Linux on IBM eServer xSeries 366 Servers


Oracle E-Business Suite 11i Payroll - Number of employees: 5,000
VendorOSHardware Config#ThreadsTime (min)Avg. CPU%Hourly Throughput
UNISYSLinux: RHEL 4 Update 3DB/App/Batch server: 1x Unisys ES7000/one Enterprise Server (4x 3.0 GHz Dual-Core Intel Xeon 7041 processors, 32 GB memory)1215.18 min53.22%57,915
IBMNovell SUSE Linux Enterprise Server 9 SP1DB, App servers: 2x IBM eServer xSeries 366 4-way server (4x 3.66 GHz Intel Xeon MP Processors (EM64T), 32 GB memory)128.42 min50+%235,644
SunSolaris 10 8/07DB/App/Batch server: 1x Sun SPARC Enterprise T5220 (1x 1.4 GHz 8-core UltraSPARC T2 processor, 64 GB memory)8 to 644.63 min40.77%64,748


Better results were highlighted. The results speak for themselves. One 1.4 GHz UltraSPARC T2 processor outperformed four 3 GHz / 3.66 GHz processors in terms of the average CPU utilization and most importantly in the hourly throughput (Hourly throughput calculation relies on the total elapsed time).

Before we conclude, let us reiterate few things purely based on the factual evidence presented in this blog post:
  • Sun's CMT servers like T2000, T5220, T5240 (two socket system with UltraSPARC T2 Plus processors) are good to run batch workloads like Oracle Applications Payroll 11i

  • Sun's CMT servers like T2000, T5220, T5240 are good to run the Oracle 10g RDBMS when the DML/DDL/SQL statements that make up the majority of the workload are not very complex, and

  • When the application is tuned appropriately, the performance of CMT processors can outperform some of the traditional processors that were touted to deliver the best single thread performance


Footnotes

1. There is a note in the UNISYS/Payroll 11i white paper that says "[...] the gains {from running increased numbers of threads} decline at higher numbers of parallel threads." This is quite contrary to what Sun observed in its Payroll 11i experiments on UltraSPARC T2 based T5220. Higher number of parallel threads (maximum: 64) improved the throughput on T5220, where as UNISYS' observation is based on their experiments with a maximum of 12 parallel threads. Moral of the story: do NOT treat all hardware alike.

2. IBM's Payroll 11i white paper has no references to the average CPU numbers. 50+% was derived from the "Figure 3: Average CPU Utilization".
________________
Technorati Tags:
 Sun |  Solaris |  CMT |  T5220 |  T2000 |  Batch Jobs |  Niagara |  Oracle |  E-Business Suite |  Payroll
Read More
Posted in | No comments

Sunday, 6 April 2008

Defending a Traffic Citation with Request for Trial by Written Declaration

Posted on 00:48 by Unknown
(The content of this blog post is relevant only for those people who live in the parts of the United States where 'Request for Trial by Written Declaration' is an option for the defendants)

Got a citation for any kind of traffic infraction? If you plan to plead guilty or to defend yourself in the traffic court, consider sending a Request for Trial by Written Declaration (aka TBD). Trial by written declaration effectively eliminates two trips to the court; and improves the chance of winning the case as long as the defendant do not exhibit any sign of admitting the guilt.

The major steps involved in requesting the trial by written declaration are as follows:
  • Read the instructions posted for the defendants who are considering 'Trial by Written Declaration'. For example, California folks have to look at the Form TR-200 for the instructions.

  • Complete the 'Request for Trial by Written Declaration' form. Again, California drivers can fill out the Form TR-205 to request the trial by written declaration.

  • Carefully draft all the facts, evidence etc., that are relevant to the citation under 'STATEMENT OF FACTS' section in the 'Request for Trial by Written Declaration' form. It is very important that you do NOT plead guilty and NOT write anything that admits the guilt of any sort implicitly or explicitly. Admitting the guilt in any form hurts the chances of winning the case.

  • It is required to include the following sentence in the 'STATEMENT OF FACTS' section.
    I declare under penalty of perjury under the laws of the State of [INSERT_YOUR_STATE_NAME_HERE] that the foregoing is true and correct.

  • Send the bail amount in the form of a check along with the 'Request for Trial by Written Declaration' form, or pay the bail amount over the web or by phone if the traffic court accepts such payments.

  • Send the filled in 'Request for Trial by Written Declaration' form along with the bail amount, if not paid already, at least 5 days prior to the due date (excluding holidays) indicated on the traffic citation. Check the instructions and the citation for the exact deadlines.

  • The 'Request for Trial by Written Declaration' form and the bail amount must be sent via 'Certified Mail' with a request for the return receipt.

  • Once the request has been sent, it may take a while for the court to review the evidence/facts submitted by the defendant and the patrol officer before they mail the decision. So just relax and wait for court's decision. That is all there is to it.

It is strongly suggested to do diligent research about various steps involved in fighting a citation involving traffic infraction.
_______________
Technorati Tags:
 Traffic Ticket |  Traffic Citation
Read More
Posted in | No comments

Thursday, 27 March 2008

Accessing the Command Line Args of a Shell Script from an Embedded AWK script

Posted on 18:45 by Unknown
Let's start with the example of the skeleton of a shell script running on Sun Solaris.
#!/bin/sh

if [ $# -lt 2 ]; then
echo "Usage: pmap_mf.sh "
exit
fi
...
pmap -x $PIDS | grep total | awk 'BEGIN { FS = " " } {print $1,$2,$3,$4,$5} {rss+=$4}
{private+=$5} END {print "Total Private mem: "private/1024" M Total RSS mem: "rss/1024" M Total
Shared mem: " (rss-private)/1024 "M **** For 10000 user load: per user memory footprint is:
"((private/1024)+(((rss-private)/1024)/NR))/10000" MB/user"}'

Notice the highlighted hard coded number for the number of users. The goal is to read the number of users value from the command line rather than editing the script every time there is a change in the number of users.

Assuming there is a shell script with an embedded AWK script like the one shown in the above example, how the awk script can access the command line arguments supplied to the top level shell script?

It is not possible to use $1, $2 etc., because $1 and $2 will hold the string tokens returned by the awk output. For example, if awk returned "total Kb 472648 242936 57768", then $1 will hold the string "total" and $2 will have "Kb".

The solution/trick is to pass the required command line arguments from the main shell script to the awk script with the -v option of awk; and then to access the arguments supplied for the awk script either by directly referencing the argument name or by using ARGC & the ARGV array notation (just like C/C++).

Note:
/usr/bin/awk on Sun Solaris doesn't accept the -v option. Use /usr/bin/nawk or /usr/xpg4/bin/awk on Solaris.

Now with all the above information, the last line of code can be modified as shown below to get rid of the hard coded numbers:
pmap -x $PIDS | grep total | nawk -v Arg1=$2 'BEGIN { FS = " " } {print $1,$2,$3,$4,$5} {rss+=$4} {private+=$5} 
END {print "Total Private mem: "private/1024" M Total RSS mem: "rss/1024" M Total Shared mem: "
(rss-private)/1024 "M **** For ", Arg1, " user load: per user memory footprint is:
"((private/1024)+(((rss-private)/1024)/NR))/Arg1" MB/user"}'

_______________
Technorati Tags:
 UNIX |  Linux |  Shell |  AWK |  Scripting
Read More
Posted in | No comments

Saturday, 22 March 2008

Oracle 10g: Upgrading to a Higher Patchset (eg., 10.2.0.1 to 10.2.0.3)

Posted on 18:10 by Unknown
This blog post outlines the steps involved in patching an existing Oracle 10g RDBMS environment to a higher patch set with an example showing the steps for upgrading from 10.2.0.1 to 10.2.0.3. For detailed instructions, check the README page that is associated with the patch{set}. Use these instructions at your own risk.
  • First of all, Oracle wouldn't let you start the database up after simply installing the latest patch set on top of an existing Oracle RDBMS environment. The following steps must be followed before it starts up successfully. Otherwise, during startup Oracle throws errors like:
    ORA-01092: ORACLE instance terminated.
    ORA-39700: database must be opened with UPGRADE option.

  • In order to install 10.2.0.3 patch set, download the patch 5337014 from metalink.oracle.com for your OS; and install it on top of 10.2.0.1. Make sure the listener and the database are down before you apply the patch.

  • Set SHARED_POOL_SIZE and JAVA_POOL_SIZE parameters to at least 150M {in the DB initialization file}, if they are either not set or the existing value is < 150M.

  • Start up the database in upgrade mode.
    SQL> STARTUP UPGRADE
    ORACLE instance started.

    Total System Global Area 1.0385E+10 bytes
    Fixed Size 2047256 bytes
    Variable Size 1241514728 bytes
    Database Buffers 9126805504 bytes
    Redo Buffers 14729216 bytes
    Database mounted.

    Note:
    If you see an error like ORA-01589: must use RESETLOGS or NORESETLOGS option for database open with STARTUP UPGRADE, clear the logs as shown below:

    SQL> ALTER DATABASE OPEN RESETLOGS: <- NORESETLOGS doesn't work here.
    SQL> SHUTDOWN

    SQL> STARTUP UPGRADE
    ORACLE instance started.

    Total System Global Area 5553258496 bytes
    Fixed Size 2038008 bytes
    Variable Size 1241515784 bytes
    Database Buffers 4294967296 bytes
    Redo Buffers 14737408 bytes
    Database mounted.
    Database opened.

  • Finally run the command line database upgrade assistant tool, dbua, to do pre upgrade checks, to upgrade the Oracle server and to perform any post upgrade activity. Here is the syntax:
    % dbua -silent -dbname $ORACLE_SID -oracleHome $ORACLE_HOME -sysDBAUserName sys
    -sysDBAPassword change_on_install -recompile_invalid_objects true

    You may have to slightly modify the above command esp. the password part to make it work in your environment. The dbua should return a status message like Database upgrade has been completed successfully, and the database is ready to use.. Otherwise review the log file (usually $ORACLE_HOME/cfgtoollogs/dbua/logs/silent*.log), make the changes as necessary and re-run the dbua tool until it succeeds.

    Using dbua is only one way of upgrading the database to the latest patch set. You can also upgrade it with the help of dbua's graphical interface or by manually executing few SQL scripts. Check the README file that is bundled with the patch for detailed steps on using the dbua GUI or the manual execution of the required SQL scripts.

____________________
 Oracle |  RDBMS |
Read More
Posted in | No comments

Sunday, 16 March 2008

Is Oracle's PeopleSoft Really a Multi-Threaded Application?

Posted on 01:51 by Unknown
Perhaps the answer to this question is irrelevant to many of the PeopleSoft end users - but it is really important for the administrators to know what kind of application they are dealing with.

Here is a snapshot of the process statistics (prstat output) for the PeopleSoft application server processes running on a Solaris 10 system:

   PID USERNAME  SIZE   RSS STATE  PRI NICE      TIME  CPU PROCESS/NLWP      
10864 psft 353M 235M sleep 60 10 0:13:42 1.5% PSAPPSRV/11
10855 psft 353M 235M sleep 2 10 0:13:55 1.5% PSAPPSRV/11
10846 psft 353M 235M sleep 3 10 0:14:04 1.5% PSAPPSRV/11
10870 psft 353M 235M sleep 0 10 0:13:50 1.5% PSAPPSRV/11
10873 psft 353M 235M sleep 1 10 0:13:57 1.4% PSAPPSRV/11
10852 psft 353M 235M sleep 0 10 0:13:57 1.4% PSAPPSRV/11
10858 psft 353M 235M sleep 60 10 0:13:47 1.3% PSAPPSRV/11
10849 psft 349M 231M cpu0 20 10 0:13:55 1.3% PSAPPSRV/11
10867 psft 353M 235M sleep 60 10 0:13:53 1.3% PSAPPSRV/11
10861 psft 349M 231M sleep 60 10 0:13:56 1.2% PSAPPSRV/11

Notice the number of LWPs (represented by NLWP in the snapshot) that are associated with each of those PSAPPSRV processes. Just by looking at the above snapshot, one may under the impression that the PSAPPSRV (PeopleSoft Application Server process) is a multi-threaded process because it appears there are 11 worker threads actively processing the user requests.

To dig a little deeper, Solaris provides the ability to check the statistics for each of the light-weight processes (LWPs) that are associated with a process (here I'm assuming that the intended audience can differentiate a process from a light-weight process). With the help of -L option of the prstat, Solaris reports the statistics for each LWP in a given process. Let's have a close look at those stats.

  PID USERNAME  SIZE   RSS STATE  PRI NICE      TIME  CPU PROCESS/LWPID
10864 psft 353M 235M cpu32 0 10 0:01:37 1.5% PSAPPSRV/1
10864 psft 353M 235M sleep 59 0 0:00:00 0.0% PSAPPSRV/11
10864 psft 353M 235M sleep 29 10 0:00:00 0.0% PSAPPSRV/10
10864 psft 353M 235M sleep 28 10 0:00:00 0.0% PSAPPSRV/9
10864 psft 353M 235M sleep 28 10 0:00:00 0.0% PSAPPSRV/8
10864 psft 353M 235M sleep 59 0 0:00:00 0.0% PSAPPSRV/7
10864 psft 353M 235M sleep 59 0 0:00:00 0.0% PSAPPSRV/6
10864 psft 353M 235M sleep 51 2 0:00:00 0.0% PSAPPSRV/5
10864 psft 353M 235M sleep 59 0 0:00:00 0.0% PSAPPSRV/4
10864 psft 353M 235M sleep 29 10 0:00:00 0.0% PSAPPSRV/3
10864 psft 353M 235M sleep 59 0 0:00:00 0.0% PSAPPSRV/2

Notice the PID in the first column. It confirms that the above snapshot is the process stats breakdown by LWPs for a given process. Now check the output under the TIME column. That column represents the cumulative execution time for the process -- LWP, in this case. Except for the LWP #1, the exec time for rest of the LWPs is zero i.e., even though the PSAPPSRV process spawned 10 more LWPs, in reality they are not doing any work at all. When I tried to find the reason {from my counterpart at Oracle Corporation} for the creation of multiple LWPs, I was told that the multiple LWPs are a side effect of loading JRE(s) into the process address space during the run-time. Also it appears the PeopleSoft application server processes (PSAPPSRV) can process only one user request (transaction) at a time. It is the limitation of the PeopleSoft Enterprise by design.

So the bottomline is: PeopleSoft application server processes (PSAPPSRV) are not multi-threaded even though they appear to be multi-threaded from the operating system perspective.

Before we conclude, make sure you understand that the discussion in this blog post applies only to the PeopleSoft application server processes, PSAPPSRV. You cannot generalize it to the whole PeopleSoft Enterprise. For example, application engine processes (PSAESRV) that run under the control of the Process Scheduler are actually multi-threaded processes. However expanding the discussion around Process Scheduler/Application Engine is beyond the scope of this blog post.

Acknowledgements:
Sanjay Goyal
________________
Technorati Tags:
 Oracle |  PeopleSoft |  Architecture
Read More
Posted in | No comments

Thursday, 6 March 2008

Windows Vista/HP Pavilion dv9000: Fixing Y! Messenger's Webcam Not Connected Error

Posted on 10:55 by Unknown
If the built-in webcam of HP Pavilion dv9000 is not functioning properly on Windows Vista with Yahoo! Messenger, one of the following steps may help fix the issue (Courtesy: HP Customer Support).

Step #1:
  • Unistall HP Pavilion Webcam software.
    Start -> Settings -> Control Panel -> Programs -> Uninstall a program.

  • Uninstall the Webcam driver from the Device Manager as follows:
    Start -> Settings -> Control Panel -> Device Manager -> Imaging devices -> HP Webcam -> Disable(by right clicking HP Webcam).

  • Restart the computer. Upon the restart, Vista will automatically load webcam driver.

If the issue persists even after these steps or if there is no entry found in "Device Manager", then download and install the updated webcam driver from the following link:

ftp://ftp.hp.com/pub/softpaq/sp34501-35000/sp34746.exe

Step #2 (Recommended):

  • Download and install Soft Paq.s from the link given below. This package installs Microsoft fixes and enhancements for the Microsoft Windows Vista Operating System, as well as providing other fixes and enhancements.

    ftp://ftp.hp.com/pub/softpaq/sp37501-38000/sp37736.exe

  • Restart the system once the installation is done.

    If the issue persists, install the following update from Microsoft.

    http://support.microsoft.com/kb/941600

  • Restart the system once the installation is done.

  • If the issue persists, follow the below steps to isolate the issue.

    Start -> Settings -> Control Panel -> Programs and Features

    In 'Programs and Features' window, look for an item called 'Tasks'. Under 'Tasks' find 'View installed updates'. Search for KB915597; if found, select it, and finally select 'Uninstall'. If prompted for the administrator password or confirmation, type the password or confirm.

  • If the issue remains, download and install the software SP35414 from the following link:

    ftp://ftp.hp.com/pub/softpaq/sp35001-35500/sp35414.exe

    Installing this software helps to Preview, Capture, Take a Snap options on the built in webcam.

  • Even if the above step fails, try:

    Start -> Settings -> Control Panel -> Administrative Tools -> Services -> Windows Image Acquisition (WIA)

    Right click on Windows Image Acquisition (WIA) and select properties. In General tab, change the start type to Automatic and press the start button. Again right click on 'Windows Image Acquisition (WIA)' and select properties. In Recovery tab set First failure, Second failure and Subsequent failure option to Restart the service. Apply the changes and restart the system.

  • If none of the above steps resolve the issue, contact HP's customer support.

____________________________
Technorati Tags:
 HP |  Pavilion |  dv9000 |  Windows |  Vista |  Yahoo |  Messenger |  Webcam |  Troubleshooting
Read More
Posted in | No comments

Monday, 3 March 2008

PeopleSoft: Fixing "msgget: No space left on device" Error on Solaris 10

Posted on 00:14 by Unknown
When high number of application server processes are configured in a single PeopleSoft domain or in multiple domains cumulative, it is very likely that the PeopleSoft server domain boot process may fail with errors like:
Booting server processes ...
exec PSSAMSRV -A -- -C psappsrv.cfg -D CS90SPV -S PSSAMSRV :
Failed.
113954.ben15!PSSAMSRV.29746.1.0: LIBTUX_CAT:681: ERROR: Failure to create message queue
113954.ben15!PSSAMSRV.29746.1.0: LIBTUX_CAT:248: ERROR: System init function failed, Uunixerr = :
msgget: No space left on device

113954.ben15!tmboot.29708.1.-2: CMDTUX_CAT:825: ERROR: Process PSSAMSRV at ben15 failed with /T
tperrno (TPEOS - operating system error)

In this particular example, the PeopleSoft Enterprise is running on a Solaris 10 system. Fortunately the error message is very clear in this case; and the failure is related to the message queues. During the domain boot up process, there is a call to msgget() to create a message queue. If the call to msgget() succeeds, it returns a non-negative integer that serves as the identifier for the newly created message queue. However in case of a failure, it returns -1 and sets the error number to EACCES, EEXIST, ENOENT or ENOSPC depending on the underlying reason.

From the above error messages it clear that the msgget() failed with the errno set to ENOSPC (No space left on device). Man page of msgget(2) has the following explanation for ENOSPC error code on Solaris:
ERRORS
The msgget() function will fail if:
...
...
ENOSPC A message queue identifier is to be created but
the system-imposed limit on the maximum number of
allowed message queue identifiers system wide
would be exceeded. See NOTES.

NOTES
...
...

The system-imposed limit on the number of message queue
identifiers is maintained on a per-project basis using the
project.max-msg-ids resource control.

Good stuff. It has enough clues to suspect the configured number for the message queue identifiers.

Prior to the release of Solaris 10, the /etc/system System V IPC tunable, msgsys:msginfo_msgmni, could be configured to control the maximum number of message queues that can be created. The default value on pre-Solaris 10 systems is 50.

With the release of Solaris 10, majority of the System V IPC tunables were obsoleted and equivalent resource controls were created for the remaining tunables, to reduce the administrative overhead. On Solaris 10 and later versions, System V IPC can be tuned on a per project basis using the newly introduced resource controls.

On Solaris 10, The resource control, project.max-msg-ids, replaced the old /etc/system tunable, msginfo_msgmni. And the default value has been raised to 128.

Now back to the failure in PeopleSoft environment. Let's first check the current value configured for project.max-msg-ids.

  • Get the project ID.
     % id -p
    uid=222227(psft) gid=2294(dba) projid=3(default)

  • Examine the project.max-msg-ids resource control for the project with ID 3, using the prctl utility.
     % prctl -n project.max-msg-ids -i project 3
    project: 3: default
    NAME PRIVILEGE VALUE FLAG ACTION RECIPIENT
    project.max-msg-ids
    privileged 128 - deny -
    system 16.8M max deny -

Alternatively run the command ipcs -q to check the number of active message queues. Note that the project with id '3' is configured to create a maximum of 128 (default) message queues. In any case, the number of active message queues from the ipcs -q output may almost match with the configured value for the project.max-msg-ids.

Since it appears the configured PeopleSoft domain(s) needs more than 128 message queues in order to bring up all the application server processes that constitute the PeopleSoft Enterprise, the solution is to increase the value for the resource control, project.max-msg-ids, to any value beyond 128. For the sake of simplicity, let's increase it to 256 (2 * default value, that is). Again prctl utility can be used to set the new value for the resource control.
  • Login as the 'root' user
     % su
    Password:

  • Increase the maximum value for the message queue identifiers to 256 using the prctl utility.
     # prctl -n project.max-msg-ids -r -v 256 -i project 3

  • Verify the new maximum value for the message queue identifiers
     # prctl -n project.max-msg-ids -i project 3
    project: 3: default
    NAME PRIVILEGE VALUE FLAG ACTION RECIPIENT
    project.max-msg-ids
    privileged 256 - deny -
    system 16.8M max deny -

That's all there is. With the above change, the PeopleSoft Enterprise should boot up at least with no Failure to create message queue .. msgget: No space left on device errors.

Before we conclude, note that the above mentioned solution is not persistent across multiple operating system reboots. To make it persistent, create a new project with projadd command. The man page for projadd(1M) has an example showing the creation of a project.
_________________________
Technorati Tags:
 Solaris |  OpenSolaris |  Oracle |  PeopleSoft |  Troubleshooting
Read More
Posted in | No comments
Newer Posts Older Posts Home
Subscribe to: Posts (Atom)

Popular Posts

  • C++: Virtual Function
    A virtual function allows derived classes to replace the implementation provided by the base class. The compiler makes sure the replacemen...
  • UNIX/Linux: File Permissions (chmod)
    A file's permissions are also known as its 'mode'; so to change them we need to use the 'chmod' command (change mode). T...
  • C/C++: Structure Vs Union
    A structure is a collection of items of different types; and each data item will have its own memory location. Where as only one item withi...
  • Blast from the Past : The Weekend Playlist #3
    The 80s contd., The 80s witnessed the rise of fine talent - so, it is only fitting to dedicate another complete playlist for the 80s. Her...
  • Achievement Award
    Got an Achievement Award/Certificate from Sun Microsystems, in recognition for my effort with Siebel Benchmark!! =:) Related post: http:...
  • Database: Oracle Server Architecture (overview)
    Oracle server consists of the following core components: 1) database(s) & 2) instance(s) 1) database consists of: 1) datafil...
  • C/C++/Java: ++ unary operator
    #include <stdio.h> int main() { int i = 5, j = 5; int total = 0; total = ++i + j++; printf("\ntotal o...
  • Solaris/C/C++: Benefit(s) of Linker (symbol) Scoping
    Introduction By default, the static linker (ld) makes all ELF symbols global in scope. This means it puts the symbols into the dynamic symbo...
  • Linux: Frozen Xwindows
    If Xwindows seem frozen, the following simple key strokes may bring back the Xserver without the need for a reboot Two ways to kill the Xwi...
  • PHP: Memory savings with mysqlnd
    mysqlnd may save memory. In the best cases, it may consume only 50% memory as that of libmysql esp. when the client application does not mod...

Categories

  • 80s music playlist
  • bandwidth iperf network solaris
  • best
  • black friday
  • breakdown database groups locality oracle pmap sga solaris
  • buy
  • deal
  • ebiz ebs hrms oracle payroll
  • emca oracle rdbms database ORA-01034
  • friday
  • Garmin
  • generic+discussion software installer
  • GPS
  • how-to solaris mmap
  • impdp ora-01089 oracle rdbms solaris tips upgrade workarounds zombie
  • Magellan
  • music
  • Navigation
  • OATS Oracle
  • Oracle Business+Intelligence Analytics Solaris SPARC T4
  • oracle database flashback FDA
  • Oracle Database RDBMS Redo Flash+Storage
  • oracle database solaris
  • oracle database solaris resource manager virtualization consolidation
  • Oracle EBS E-Business+Suite SPARC SuperCluster Optimized+Solution
  • Oracle EBS E-Business+Suite Workaround Tip
  • oracle lob bfile blob securefile rdbms database tips performance clob
  • oracle obiee analytics presentation+services
  • Oracle OID LDAP ADS
  • Oracle OID LDAP SPARC T5 T5-2 Benchmark
  • oracle pls-00201 dbms_system
  • oracle siebel CRM SCBroker load+balancing
  • Oracle Siebel Sun SPARC T4 Benchmark
  • Oracle Siebel Sun SPARC T5 Benchmark T5-2
  • Oracle Solaris
  • Oracle Solaris Database RDBMS Redo Flash F40 AWR
  • oracle solaris rpc statd RPC troubleshooting
  • oracle solaris svm solaris+volume+manager
  • Oracle Solaris Tips
  • oracle+solaris
  • RDC
  • sale
  • Smartphone Samsung Galaxy S2 Phone+Shutter Tip Android ICS
  • solaris oracle database fmw weblogic java dfw
  • SuperCluster Oracle Database RDBMS RAC Solaris Zones
  • tee
  • thanksgiving sale
  • tips
  • TomTom
  • windows

Blog Archive

  • ▼  2013 (16)
    • ▼  December (3)
      • Blast from the Past : The Weekend Playlist #3
      • Measuring Network Bandwidth Using iperf
      • Blast from the Past : The Weekend Playlist #2
    • ►  November (2)
    • ►  October (1)
    • ►  September (1)
    • ►  August (1)
    • ►  July (1)
    • ►  June (1)
    • ►  May (1)
    • ►  April (1)
    • ►  March (1)
    • ►  February (2)
    • ►  January (1)
  • ►  2012 (14)
    • ►  December (1)
    • ►  November (1)
    • ►  October (1)
    • ►  September (1)
    • ►  August (1)
    • ►  July (1)
    • ►  June (2)
    • ►  May (1)
    • ►  April (1)
    • ►  March (1)
    • ►  February (1)
    • ►  January (2)
  • ►  2011 (15)
    • ►  December (2)
    • ►  November (1)
    • ►  October (2)
    • ►  September (1)
    • ►  August (2)
    • ►  July (1)
    • ►  May (2)
    • ►  April (1)
    • ►  March (1)
    • ►  February (1)
    • ►  January (1)
  • ►  2010 (19)
    • ►  December (3)
    • ►  November (1)
    • ►  October (2)
    • ►  September (1)
    • ►  August (1)
    • ►  July (1)
    • ►  June (1)
    • ►  May (5)
    • ►  April (1)
    • ►  March (1)
    • ►  February (1)
    • ►  January (1)
  • ►  2009 (25)
    • ►  December (1)
    • ►  November (2)
    • ►  October (1)
    • ►  September (1)
    • ►  August (2)
    • ►  July (2)
    • ►  June (1)
    • ►  May (2)
    • ►  April (3)
    • ►  March (1)
    • ►  February (5)
    • ►  January (4)
  • ►  2008 (34)
    • ►  December (2)
    • ►  November (2)
    • ►  October (2)
    • ►  September (1)
    • ►  August (4)
    • ►  July (2)
    • ►  June (3)
    • ►  May (3)
    • ►  April (2)
    • ►  March (5)
    • ►  February (4)
    • ►  January (4)
  • ►  2007 (33)
    • ►  December (2)
    • ►  November (4)
    • ►  October (2)
    • ►  September (5)
    • ►  August (3)
    • ►  June (2)
    • ►  May (3)
    • ►  April (5)
    • ►  March (3)
    • ►  February (1)
    • ►  January (3)
  • ►  2006 (40)
    • ►  December (2)
    • ►  November (6)
    • ►  October (2)
    • ►  September (2)
    • ►  August (1)
    • ►  July (2)
    • ►  June (2)
    • ►  May (4)
    • ►  April (5)
    • ►  March (5)
    • ►  February (3)
    • ►  January (6)
  • ►  2005 (72)
    • ►  December (5)
    • ►  November (2)
    • ►  October (6)
    • ►  September (5)
    • ►  August (5)
    • ►  July (10)
    • ►  June (8)
    • ►  May (9)
    • ►  April (6)
    • ►  March (6)
    • ►  February (5)
    • ►  January (5)
  • ►  2004 (36)
    • ►  December (1)
    • ►  November (5)
    • ►  October (12)
    • ►  September (18)
Powered by Blogger.

About Me

Unknown
View my complete profile