Tuesday, October 18, 2016

JDK 8 - HOTSPOT JVM


With JDK 8 getting adopted vastly, I just want to give a high level view of hotspot memory management and the garbage collection techniques


Quickly, about Garbage Collection 
Allocates objects to young generation
Promotes aged objects to old generation
Marking old generation objects
Recovering space by removing unreachable or compacting live objects


Available collectors in Hotspot JVM
1.       Serial collector – for single processor and low concurrent apps. -XX:+UseSerialGC.

2.       Parallel collector (or throughput collector) – by default one on the server machines and medium to large scale apps. User -XX:+UseParallelGC. Its multithreaded GC. Select this when high throughput is needed.

3.       Mostly concurrent collector – performs most of its work concurrently i.e. app is still running. Select this when faster response times are needed. Since the app is running most of the time without any pauses, means RTs are not impacted often by the pause times. There are 2 types of mostly concurrent collectors – Concurrent Mark & sweep, G1 (garbage first gc – a generational algo preferred for heaps more than 4GB).

*throughput is termed as time spent in GC vs time spent in app. The more timespent is app means it has high amount of resources to deliver. 
* for generational GCs i.e. GCs with young,eden,tenured generations, both young and major gcs do pause the app in all types of GC algos. Minor/young gc happens very fast but a major gc happens when there is no space for tenured objects, so it involves entire heap collections i.e. identifying the reachable objects, marking and sweeping unreachable objs and compacting the space to avoid fragmentation. So it pauses longer time. CMS, G1 tries to reduce this by doing them always in the background while app is running and pausing app only in certain times and for short duration

Tuning first steps:
Go with the defaults
-          serial for single core machine,
-          parallel for high throughput requirement i.e. where App’s RTs can be acceptable at a bit higher sec that means giving longer pause times i.e. less frequent pauses – overall, the app will be paused at times but longer (means high throughput but breaches RT as it causes spike) and the parallel GC does its work with multiple threads in parallel to each and consumes CPU.
-          Mostly concurrent – on large apps, multi processor machines, when RT is important. Since RT is important, the pause times are less but may be frequent however, some part of the gc work like marking can happen without any pause. Use CMS when heap  when heap less than 4gb else use G1
-          Once the default is selected, check whether the requirement is met else, try playing around increasing the XMX. Then try tuning the heap areas.


Parallel Collector:
-          -XX:+UseParallelGC
-          both minor and major collections are executed in parallel
-          number of threads is calculated as 5/8 * N (number of HW threads) or can be set manually as -XX:ParallelGCThreads=
-          since higher number of threads causes fragmentation when promoting objects from young to tenured spaces, try reducing the number or increase tenured space to cope up with freagmentation.
-          parallel collector tunes automatically based on behavior. So no need to specify generation sizes or granular level tunings. Behavior is mentioned are max gc pause time, throughput, foot print of heap size.
-          Max pausetime target:          -XX:MaxGCPauseMillis=
-          Throughput target: -XX:GCTimeRatio=  i.e. it tries to spend 1/(1+N) percentage in GC. By default is 99 i.e. i/(1+99) = 1% in GC , 99% leaving it to app
-          Footprint : is basically the –Xmx
-          parallel collector tries to meet max pausetime first then throughput then footprint targets. It plays around growth/shrink percentages of generations to achieve these targets.

Mostly concurrent collectors:
CMS: if the importance is faster RT i.e. lesser pause times and can afford CPU
-          -XX:+UseConcMarkSweepGC.
-          reduces pause time by always keeping few threads to concurrently (i.e. without pausing the threads) to mark the reachable & unreachable objects and sweeping of unreachable objects but pauses during a major gc during movement of references. It always tries to keep the tenured space clean to avoid longer pauses or accumulation of objects. When it fails keep up the tenured space free then a full collection happens with all the threads paused (a failure case)
-          In CMS, the tenured and young can happen independently as they have different threads running in concurrent to the application.
-          CMS output is a bit different to other GC outputs. -verbose:gc and -XX:+PrintGCDetails (-XX:+PrintAllGCDetails to print more details). It has below output
       CMS-initial-mark indicates the start of the concurrent collection cycle
                                 CMS-concurrent-mark indicates the end of the concurrent marking phase
CMS-concurrent-preclean
                CMS-remark
                                                                                CMS-concurrent-sweep marks the end of the conc sweeping phase
                                CMS-concurrent-reset (getting ready for next collection)
-          CMS and Parallel does compactions for the whole heap when there is no consecutive space available

G1: for large heaps and pause times targets can be met at higher probability and also achieving high throughput (i.e. more time given to app)
-       heap will be partitioned into a set of equally sized heap regions, each a contiguous range of virtual memory. the algo performs a concurrent global marking phase to determine the live objects of the heap. After the marking phase completes, it collects the mostly empty regions first, to yield a large amount of free space. So it is called Garbage-First. 
-          This always works to reduce fragmentation by compacting during collections.
-          G1 is beneficial when there is large amount of live data i.e. 50% of heap, allocation rate changes, when there is long collection time or if there is long compaction times

Default configuration on server class machines
On server-class machines, the following are selected by default if not specified otherwise.
Throughput garbage collector
Initial heap size of 1/64 of physical memory up to 1 GB
Maximum heap size of 1/4 of physical memory up to 1 GB
Server runtime compiler


And, to view the default configurations use java -XX:+PrintFlagsFinal -version  


Sunday, October 2, 2016

JMeter Solr Banana


Want to give some colors to JMeter ? We all know JMeter is a great tool and helps load testing the apps.. how about bringing two more awesome tools on to the table - Solr and Banana

With Solr, we can store/index/search large amount of data and Banana is a pretty tool with lot of latest HTML and javascript capabilities to draft some cool graphs to show the trends.

So, how we can use these tools ..

we can do a lot in fact, but to start with,

- Load test results to display the runtime stats - Response times/Transaction throughput/Bytes and the list goes on..
- Monitoring system resources like Memory/CPU/Disk/Load, JVM garbage collection activity etc.,

So, its all sounds interesting ? and do you think we can really build a good monitoring tool ? Well,  below are couple of dashboards

Load Test Report:



System Resources


These are few sample dashboards. The actual setup does a lot more.. the remote agents collect the data and pushes them to solr and the dashboard refreshes the stats.



Saturday, September 24, 2016

Sizing JVMs and VM Memory



How to estimate the amount of JVM heap and host memory requirement in complex cases where there are more JVM instances per VM and more apps per JVM and have different usage requirement per app ?

let's see some of the metrics to collect to construct an equation on estimating the sizes.


  • list the apps per instance. Ideally there will be multiple apps deployed on one instance of JVM and all these apps are not equally used and each app will have its own characteristics thus have uneven memory requirements i.e. based on req. classes, code logic and constructs etc.,
  • And identify the 'typical' transactions.. ideally, in any app, 20% of features are executed 80% of times
  • Take a first measurement - server start up heap. Once the server is up completely, get the first point after a full GC (hotspot) or first OC (jrockit)
  • Now, ideally the first user access is demanding one as it loads lot of stuff. So, do just logins into each of the app and take a note of OCs/FGCs for each user. So, per app, the numbers gives the amount of heap required for first login
  • So, XMS - to make it good enough, it can be sum of (server startup heap + sum of all above deltas )
  • Now, exercise the typical 20% transaction as single user per app and do not logout - so these are typical active users on the system - call these deltas as h1, h2 etc . This can be on a warmed up server but best is to have no other logged in users in the system. Each delta can be noted or take the final delta after all different app users are in and executed the typical flows but not logged out.
  • XMX i.e. the max amount of memory can be XMS + ( total delta * (%app1 conc. users)   +  ... )

for hotspot jvms or other jvms where there are more heap spaces like young/survivor, perm etc., they can be calculated as proportionate ratio of XMX ..

And, typical host memory usage requirement can be calculated as approx 1.8 times of XMX

If someone can simulate load and want to bare the cost, time and complexity of all setup to do that then it is probably best to get the numbers based on test results.., the benefits in doing so can be not just limited to JVM heap but gives the other resource pools across the layers..

Wednesday, September 14, 2016

Mobiles - The way they transformed made things reachable to some but created complexity to others


I am talking about the quality checks for the applications built for mobile devices. Mobile devices transformation has to be considered as the fastest change in technology so as the complexity in delivering applications on this new platform which is becoming a must !

Besides making sure the features work i.e. testing the functionality, it is the most important task to ensure the app actually perform as expected. someone can say what is there to think about as it is a light weight code .. there is infact a lot that could disrupt the performance - examples - have to support different vendors like android/APPLE, version of operating systems, underlying hardware, resolutions, GPU, CPUs, memory, network carriers and subscribed bandwidths like 2G/3G/4G, inter-apps interruptions, backend thread support for async calls .. what not !

Well there is a lot, I can think of how one can approach to validate performance of the mobile applications and the upstream servers by using some of the great tools in the market.

Below are some of the tasks to evaluate performance and what these tools can do..


  • Test and monitor the app's performance on different emulators. If it has to be done independently then each IDE is needed like android studio, IOS IDE and respective skill set as well. Perfecto addresses this with its support to multiple devices 
  • Test the app's performance on different real devices- can't do it manually.. can not have a farm of devices.. perfecto can do it using its mobile device cloud.. or its emulator is good enough as well
  • Simulate apps performance under different network speeds on devices/emulators - can do by using various softwares like act but again perfecto along with shunra integration can manage devices by emulating various network bandwidths
  • Capture the traffic and simulate 1000's of devices load to app servers - this can be done using emulators and tcp capture.. but perfecto-shunra can do and the captured traffic can be fed to Loadrunner. And, Shunra can virtualize networks to emulate load from different locations/carriers/bandwidths during load tests
  • At the same time one would have to recheck the app's performance on the device while server under load  - again by perfecto

Its a different world now with mobiles/tablets/wearables that move .. Apps need to support all these !
Gone are those days where apps are accessed from a standalone PC ..



Wednesday, August 24, 2016

End to end request processing time


In a simplistic view, in a typical online system, this is where one needs to check to know any slowness




Browser rendering time - if the page size too big or too complex with java scripts and style sheets
then the rendering time could take time..
# of static content being used on the web page has an impact on the overall load time. And, if they are not cacheable, the round trips even from a CDN could have an impact on the load time
Total download time for a page depends on how big is the response content and how many susbrequests are triggered part of the page and the network time and entire server side processing time
Page load time is the time by when user sees the page.. so above all have impact on it.

coming to network, it is important to know what is the path the request is traversing from client to the server. Is it taking longest path via CDN or how the addresses being resolved over public internet and any proxy being used etc., the network delay and packet loss are important factors to keep an eye on.

coming to server side, the request processing time depends on many factors.. it depends on underlying infrastructure, resources, architecture, code logic implementations etc.,
but there are few check points to look at to break it down..
checking at http server layer tells the variation between end page load time and total server time. This helps identify any network delays.
checking the difference between http server to app server times tells if there is any delay in the middle layers like authentication routing..
on the app server side, it could spend time in many places.. the routing can happen to many servers or even to external systems. using runtime instrumentation tools, it is possible to break down the time spent in pure code, wait times due to synchronous code blocking, gc pause times, time spent in reading from sockets/wirte to while interacting with DB or in the calls to other servers while making remote calls like rjvm to EJBs or service calls etc., by breaking this way, each underlying activity and the delays can be identified

Each of the above ones are not more than just an index to an ocean of tuneable metrics that each underlying technology modules contains. However, where to look at and what to tune is the key.



Application performance - odd bits



we do normally warm up the systems and concerned about the how it performs under load but will it meet customer expectation or improve customer experience or avoid frustrating scenarios ?
how about the 'first time' and idle case performances ?

did you ever check what is the single user response times and request processing times on server side ? if it is not performing any good for single user then it will not do for multiple. The base performance for single user is what needs to tuned first

what about first time cases ? we can not ask users of the system to hang on until system warms up ! then who will use it first :)  so, analysing cold case system performance is important. it is important to know what happens on first access across the layers in cold cases i.e. after a system restart, app servers or entire VMs or DB etc., and what needs to be tweaked to get better first time performance.

how about sleeping systems ! not all systems work round the clock atleast not all days in a week.. so did you ever check what is the resource usage of idle systems ? there could be some unexpected code path executions might happen even under no load case. So it is important to know the resources usage under no load as they might also run all the time.

how about apps performance under a very slow network access or networks with high packet loss. how reliable are the systems..

what happens when infrastructure fails ? how many users can still continue doing what ever they are doing with out any interruption and feel the same performance of the app while a DR process kicks in..

how about the systems with long running sessions.. users may not logout for long time and it is required to keep the sessions so long and what is the impact !

how the applications handle when majority of end users do not logout but just closes the browser ! how to handle the memory usage in those cases..



Saturday, August 20, 2016

Java EE - EJB


Enterprise Java Beans - EJB

EJBs are the Java EE server side components which implement app's business logic. They normally be deployed in EJB containers provided by the app servers. Implementing EJBs can provide scalability to the application and better handling of security and transactions. EJBs can also implement webservices.

Types of EJB:
Session - To implement user actions
- stateful: maintains a state for the client and this bean can not be shared.
- stateless: does not maintain state so can be allocated to any clients. They are reusable. So this
        pool will be less compared to stateful.
- singletone session beans: have only one instance for the whole time.  They can get initiated when the app starts. Although they act as stateless but there is no pool because of single existence

  Message Driven Beans (MDB): for listening to the messages either from queues or JMS
(if someone read about entity beans, they are now part of persistence API..(referring to Java EE7)

Implementation is simple.. infact, the annotations made it so..lets says to write a simple stateless EJB, just annotate the class with @Stateless and implement the business logic.. and to call this EJB, a servlet can do ..ofcourse, you need to annotate with @WebServlet(urlPatterns="/") the url pattern says the context root. And, as usual extend the class with HttpServlet. And, to access the EJB, just annotate with @EJB and then the declare the EJB instance.

However, in a typical EJB implementation, there could be all type of session beans - stateless, stateless bean implementing a service, stateful bean accessed remotely etc.,

A remote interface is required for the beans which allow remote access. This remote business interface defines all the business methods of a bean and annotated with Remote from javax.ejb pacakge and it gets implemented by the session bean. A session bean can be an end point for a web service.

A stateful session bean can have methods annotated with Remove which can be invoked by client to remove the instance.

In the case of singleton bean, the concurrent access from the clients can be controlled in two ways - container managed or bean managed by annotating accordingly. And, the methods must be annotated with the locktype i.e. read or write so that concurrent access can allowed or provided with a synchronous mechanism respectively..

For stateless session beans to implement service end points, they must be annotated with @Webservice and the business methods that are exposed must be annotated with @Webmethods. They can also implement async methods so that clients no need to wait for response from long running methods.

Coming to EJB pools - in weblogic, there is an element called max-beans-in-free-pool in weblogic-ejb-jar.xml. This determines how many EJBs must be made available in free pool. max-beans-in-pool will put a cap on the pool limit. For MDBs, the container will create as many instances required based on the size limited to max-beans-in-free-pool. Default MDB threads are 16 but this can be changed by having custom queue or workmanagers