In a previous post, I mentioned a favourite interview question of mine: Stored Process Server vs Workspace Server - What's the difference? I thought I'd mention another. Outside of incidental information picked-up through your regular work, how do you gain new knowledge?
If you rely on your regular work tasks to guide you into learning new stuff, you're likely to be limited to a slow pace of learning, and likely to learn more about what you already know, but much less likely to learn completely new concepts and technologies. For instance, you don't *have* to use hash tables, there are plenty of adequate alternatives, but once you've got to grips with the concepts of hash tables you'll see how you might approach your work in a different way and/or approach tasks that you'd previously thought too difficult.
Take regular expressions as another example. SAS has a raft of string functions, so why learn a completely new way of handling strings? Well, those who've put time into understanding regular expressions will tell you that the ability to do complex string functions in a simple fashion with regular expressions makes your coding and testing simpler too.
And, it doesn't have to be code syntax that you learn. Perhaps you could make better use of Enterprise Guide's interface and features, strengthen your adoption of modular design, or gain a stronger grasp of the concepts and rationale for change management.
So, voluntarily learning new things (and practising their use with code katas, maybe) is a "good thing". How might you go about it?
There are two broad approaches. The first is to pick topics of your own choosing, and then go and research them; the second is to be randomly given new topics to understand and learn.
For the first approach, you have a choice of courses and books, but you also have a cornucopia of information on the world wide web which is available through a small number of well-chosen searches. Conference papers are an excellent, readable source of information.
Taking the latter approach, you might subscribe to one or more blogs such as NOTE:. Not every article will be of interest or relevance to you, but you shouldn't feel obliged to read and learn everything that comes your way! My December article on the use of blogs and RSS readers spelt out how to make the best use of tools such as Google Reader, and offered advice on the discipline required to read when you've got time, and discard when you don't.
Attending conferences offers opportunities for both approaches. You might attend papers because they are focused on a topic of specific interest, but you might also choose sessions from the agenda in a less focused approach which is driven more by a sense of curiosity.
Having encouraged you to learn new things (and I do strongly encourage everybody to do so), I should offer a word of caution. Code has to be maintained and supported; by people other than yourself. So it is unforgivable to write code that is so complex that it is beyond a reasonable expectation of understanding by others. When you do learn something new and distinctly different (perhaps hash tables are an example), always take the opportunity to share your new-found knowledge with others. Your team mates will appreciate it and respect you all the more for doing so.
What steps to new learning will you take today?
SAS® and software development best practice. Hints, tips, & experience of interest to a wide range of SAS practitioners. Published by Andrew Ratcliffe's RTSL.eu, guiding clients to knowledge since 1993
Tuesday, 5 March 2013
Wednesday, 27 February 2013
NOTE: DS2 Final, Final Comments
I got a lot of feedback about the DS2 articles I recently wrote. I should add a few more summarising comments:
DS2:
NOTE: DS2. Data Step Evolved?
NOTE: DS2, Learn Something New!
NOTE: DS2, SQL Within a SET Statement
NOTE: DS2, Threaded Processing
NOTE: DS2, Final Comments
NOTE: DS2, Final, Final Comments
- I mentioned that DS2 doesn't currently support MERGE. Chris Hemedinger commented that the ability to use SQL in your SET statement means that you can use a SQL join on the SET statement and thereby a) achieve a MERGE, and b) benefit from "pushing" the work to a source database (if using non-SAS data).
- Jack Hamilton observed that accidentally omitting the QUIT at the end of a call to PROC DS2 will cause some less-than-obvious syntax errors for the following conventional DATA step. Beware!
- With power comes great responsibility! Whilst a DS2 program and its executing threads are constrained to one process on one machine, there are currently no limits to how many threads a user can request to use with thread programs. So, a user could take full advantage of all the cores on the machine - to the detriment of their fellow users.
DS2 takes no account of the CPUCOUNT option, and it has no association with SAS Grid controls (with which, SAS administrators are able to manage and control users' use of system resources). I'm sure that some degree of control, facilitating workload balancing, will follow in time
DS2:
NOTE: DS2. Data Step Evolved?
NOTE: DS2, Learn Something New!
NOTE: DS2, SQL Within a SET Statement
NOTE: DS2, Threaded Processing
NOTE: DS2, Final Comments
NOTE: DS2, Final, Final Comments
Tuesday, 26 February 2013
NOTE: Plan Early and Avoid a Year 2020 Problem!
My friend Stuart Pearson alerted me to a distant but fast approaching issue for some SAS sites. Stuart observed a problem at one of his clients as they were working on forecasts for the next ten years and one of their processes was feeding-in 2 digit years. 2021 was becoming 1921. This was because the default YEARCUTOFF option value is 1920. As Stuart said to me, as we near 2020 it is something that SAS users need to be fully aware of.
It's not a difficult issue to overcome. But to overcome it, you need to be aware of it.
Stuart also highlighted that SAS Data Quality handles 2 digit years differently to SAS/BASE. By default, anything greater than 04 and less than 99 will be assigned to the 20th century (19nn); for years 01 - 03, the 21st century (20nn) is assigned. You can create your own schemes and rules, but you need to be aware of the defaults.
It's not a difficult issue to overcome. But to overcome it, you need to be aware of it.
Stuart also highlighted that SAS Data Quality handles 2 digit years differently to SAS/BASE. By default, anything greater than 04 and less than 99 will be assigned to the 20th century (19nn); for years 01 - 03, the 21st century (20nn) is assigned. You can create your own schemes and rules, but you need to be aware of the defaults.
Monday, 25 February 2013
Giving Focus to Peer Reviews
I'm a keen advocate for peer reviews. I've written about them before (here and here) but there's always more to say.
Peer reviews must always be treated as a constructive exercise. The style of questions can play a big part in the atmosphere and tone of the exercise.
There are many ways you can judge the quality of code. It's easy for developers to put too much time and effort in the wrong places, building things no one would use to begin with. So, aside from the basic questions of the code meeting design and coding guidelines, I ask myself the following questions:
a. How easy will it be to add new features?
b. How easy will it be to change existing features?
c. How easy will it be for a new team member to become productive?
I look for a good balance between the potentially contradictory questions. Using complex architecture might help with (1) and sometimes (2) but will probably hurt (3), so there is an interesting compromise to be delivered.
These questions ensure that the review takes account of the productivity of the team (and company) in addition to the regular technical factors.
Peer reviews must always be treated as a constructive exercise. The style of questions can play a big part in the atmosphere and tone of the exercise.
There are many ways you can judge the quality of code. It's easy for developers to put too much time and effort in the wrong places, building things no one would use to begin with. So, aside from the basic questions of the code meeting design and coding guidelines, I ask myself the following questions:
a. How easy will it be to add new features?
b. How easy will it be to change existing features?
c. How easy will it be for a new team member to become productive?
I look for a good balance between the potentially contradictory questions. Using complex architecture might help with (1) and sometimes (2) but will probably hurt (3), so there is an interesting compromise to be delivered.
These questions ensure that the review takes account of the productivity of the team (and company) in addition to the regular technical factors.
Wednesday, 20 February 2013
NOTE: DS2, Final Comments #sasgf13
In my previous posts, I've covered many aspects of DS2 (previous posts are listed at the bottom of this post). It's time to wrap up by offering a few more final details.
Whilst DS2 provides a wide range of data types, not all types are supported by all output data structures. The BASE engine, for example, has not been updated to allow storage of anything other than numeric and character variables in SAS datasets, so an attempt to create a data set with a variable of type BIGINT will be met with a warning message:
WARNING: BASE driver, creation of a BIGINT column has been requested, but is not supported by the BASE driver. A DOUBLE PRECISION column has been created instead.
The traditional SAS numeric variable is known in this context as a double precision column!
Paired with the SAS Embedded Process, DS2 enables you to perform processing similar to SAS in completely new places, such as in-database processing in relational databases, the SAS High-Performance Analytics grid and the DataFlux Federation Server.
If you want to know more, consider attending Mark Jordan's pre-conference tutorial at this year's SAS Global Forum. In the seminar, entitled "What Will DS2 Do for You?", you will learn the basics of writing and executing DS2 code. Mark promises to shows attendees how to explicitly control threading to leverage multiple processors when performing data manipulation and data modeling. He will demonstrate how DS2 improves extensibility and data abstraction in your code through the implementation of packages and methods. It's an extra fee event ($155) but could add a powerful extra string to your SAS bow!
DS2:
NOTE: DS2. Data Step Evolved?
NOTE: DS2, Learn Something New!
NOTE: DS2, SQL Within a SET Statement
NOTE: DS2, Threaded Processing
NOTE: DS2, Final Comments
NOTE: DS2, Final, Final Comments
Whilst DS2 provides a wide range of data types, not all types are supported by all output data structures. The BASE engine, for example, has not been updated to allow storage of anything other than numeric and character variables in SAS datasets, so an attempt to create a data set with a variable of type BIGINT will be met with a warning message:
WARNING: BASE driver, creation of a BIGINT column has been requested, but is not supported by the BASE driver. A DOUBLE PRECISION column has been created instead.
The traditional SAS numeric variable is known in this context as a double precision column!
Paired with the SAS Embedded Process, DS2 enables you to perform processing similar to SAS in completely new places, such as in-database processing in relational databases, the SAS High-Performance Analytics grid and the DataFlux Federation Server.
If you want to know more, consider attending Mark Jordan's pre-conference tutorial at this year's SAS Global Forum. In the seminar, entitled "What Will DS2 Do for You?", you will learn the basics of writing and executing DS2 code. Mark promises to shows attendees how to explicitly control threading to leverage multiple processors when performing data manipulation and data modeling. He will demonstrate how DS2 improves extensibility and data abstraction in your code through the implementation of packages and methods. It's an extra fee event ($155) but could add a powerful extra string to your SAS bow!
DS2:
NOTE: DS2. Data Step Evolved?
NOTE: DS2, Learn Something New!
NOTE: DS2, SQL Within a SET Statement
NOTE: DS2, Threaded Processing
NOTE: DS2, Final Comments
NOTE: DS2, Final, Final Comments
Labels:
Performance,
SAS,
SGF,
Syntax,
Training
Monday, 18 February 2013
NOTE: DS2, Threaded Processing
In my recent posts on DS2 (DATA step evolved), I showed the basic syntax plus packages & methods, and I showed the use of SQL within a SET statement. In today's post, I'll show the biggest raison d'être for DS2 - the ability to run your code in threads to make it finish its job more quickly.
"Big data", that's the big talking point. One of the key principles of performing speedy analytics on big data is to split the data across multiple processors and disks, to send the code to the distributed processors and disks, have the code run on each processor against its sub-set of data, and to collate the results back at the point from which the request was originally made. Thus, we're sending code to the data rather than pulling the data to the code. It's quicker to send a few dozen lines of code to many processors than it is to pull many millions of rows of data to one (big) processor.
DS2 was designed for data manipulation and data modeling applications. DS2 also enhances a SAS programmer’s repertoire with object-based tools by providing data abstraction using packages and methods. DS2 executes both within a SAS session by using PROC DS2, and within selected databases where the SAS Embedded Process is installed.
Here's a simple example. In summary, it shows how the use of eight threads reduces the turnaround time of the task from 24 seconds in a conventional DATA step to 4.4 seconds in a call to DS2 with eight threads.
SAS (r) Proprietary Software Release 9.2 TS2M3
CPUCOUNT=24 Number of processors available. 29
30 /****************************/
31 /* Create a chumky data set */
32 /****************************/
33 data work.jmaster;
34 do j = 1 to 10e6;
35 output;
36 end;
37 run; 38
39 /**************************/
40 /* Now read it three ways */
41 /**************************/
42
43 /* But first define the threaded code thread */
44 proc ds2;
45 thread r /overwrite=yes;
46 dcl double count;
47 method run();
48 set work.jmaster;
49 count+1;
50 do k=1 to 100;/* Add some gratuitous computation! */
51 x=k/count + k/count + k/count;
52 end;
53 end;
54 method term();
55 OUTPUT;
56 end;
57 endthread;
58 run;
NOTE: Execution succeeded. No rows affected.59 quit; 60
61 /* One thread */
62 proc ds2;
63 data j1(overwrite=yes);
64 dcl thread r r_instance;
65 dcl double count;
66 method run();
67 set from r_instance threads=1;
68 total+count;
69 end;
70 enddata;
71 run;NOTE: Execution succeeded. One row affected.72 quit;NOTE: PROCEDURE DS2 used (Total process time):
real time 25.09 seconds
cpu time 25.16 seconds73
74 /* Eight threads */
75 proc ds2;
76 data j8(overwrite=yes);
77 dcl thread r r_instance;
78 dcl double count;
79 method run();
80 set from r_instance threads=8;
81 total+count;
82 end;
83 enddata;
84 run;NOTE: Execution succeeded. 8 rows affected.85 quit;NOTE: PROCEDURE DS2 used (Total process time):
real time 4.40 seconds
cpu time 32.96 seconds86
87 /* And read it in DATA step */
88 data jold;
89 set work.jmaster end=finish;
90 count+1;
91 do k=1 to 100;/* Add some gratuitous computation! */
92 x=k/count + k/count + k/count;
93 end;
94 if finish then output;
95 run;NOTE: There were 10000000 observations read from the data set WORK.JMASTER.NOTE: The data set WORK.JOLD has 1 observations and 4 variables.
NOTE: DATA statement used (Total process time):
real time 23.98 seconds
cpu time 23.75 seconds
The code does six things:
This example suggests that great benefit can be achieved with DS2 threads. In our simple case, the code was compute-bound. If your task is more I/O-bound then the benefits may be less predictable. Jason, from the DS2 development team, recently told me:
In my next post, I'll wrap up the topic with a few extra details.
DS2:
NOTE: DS2. Data Step Evolved?
NOTE: DS2, Learn Something New!
NOTE: DS2, SQL Within a SET Statement
NOTE: DS2, Threaded Processing
NOTE: DS2, Final Comments
NOTE: DS2, Final, Final Comments
"Big data", that's the big talking point. One of the key principles of performing speedy analytics on big data is to split the data across multiple processors and disks, to send the code to the distributed processors and disks, have the code run on each processor against its sub-set of data, and to collate the results back at the point from which the request was originally made. Thus, we're sending code to the data rather than pulling the data to the code. It's quicker to send a few dozen lines of code to many processors than it is to pull many millions of rows of data to one (big) processor.
DS2 was designed for data manipulation and data modeling applications. DS2 also enhances a SAS programmer’s repertoire with object-based tools by providing data abstraction using packages and methods. DS2 executes both within a SAS session by using PROC DS2, and within selected databases where the SAS Embedded Process is installed.
Here's a simple example. In summary, it shows how the use of eight threads reduces the turnaround time of the task from 24 seconds in a conventional DATA step to 4.4 seconds in a call to DS2 with eight threads.
16 /*****************************/
17 /* Create a chunky data set. */
18 /* Then read it: */
19 /* a. With one thread */
20 /* b. With eight threads */
21 /* c. Using "old" DATA step */
22 /*****************************/
23
24 options msglevel=n;
25 options cpucount=actual;
26
27 proc options option=threads;run;SAS (r) Proprietary Software Release 9.2 TS2M3
THREADS Threads are available for use with features of the SAS System that support threading 28 proc options option=cpucount;run;SAS (r) Proprietary Software Release 9.2 TS2M3
CPUCOUNT=24 Number of processors available. 29
30 /****************************/
31 /* Create a chumky data set */
32 /****************************/
33 data work.jmaster;
34 do j = 1 to 10e6;
35 output;
36 end;
37 run; 38
39 /**************************/
40 /* Now read it three ways */
41 /**************************/
42
43 /* But first define the threaded code thread */
44 proc ds2;
45 thread r /overwrite=yes;
46 dcl double count;
47 method run();
48 set work.jmaster;
49 count+1;
50 do k=1 to 100;/* Add some gratuitous computation! */
51 x=k/count + k/count + k/count;
52 end;
53 end;
54 method term();
55 OUTPUT;
56 end;
57 endthread;
58 run;
NOTE: Execution succeeded. No rows affected.59 quit; 60
61 /* One thread */
62 proc ds2;
63 data j1(overwrite=yes);
64 dcl thread r r_instance;
65 dcl double count;
66 method run();
67 set from r_instance threads=1;
68 total+count;
69 end;
70 enddata;
71 run;NOTE: Execution succeeded. One row affected.72 quit;NOTE: PROCEDURE DS2 used (Total process time):
real time 25.09 seconds
cpu time 25.16 seconds73
74 /* Eight threads */
75 proc ds2;
76 data j8(overwrite=yes);
77 dcl thread r r_instance;
78 dcl double count;
79 method run();
80 set from r_instance threads=8;
81 total+count;
82 end;
83 enddata;
84 run;NOTE: Execution succeeded. 8 rows affected.85 quit;NOTE: PROCEDURE DS2 used (Total process time):
real time 4.40 seconds
cpu time 32.96 seconds86
87 /* And read it in DATA step */
88 data jold;
89 set work.jmaster end=finish;
90 count+1;
91 do k=1 to 100;/* Add some gratuitous computation! */
92 x=k/count + k/count + k/count;
93 end;
94 if finish then output;
95 run;NOTE: There were 10000000 observations read from the data set WORK.JMASTER.NOTE: The data set WORK.JOLD has 1 observations and 4 variables.
NOTE: DATA statement used (Total process time):
real time 23.98 seconds
cpu time 23.75 seconds
The code does six things:
- Firstly, it sets the maximum number of CPUs available to the SAS task. It sets this to 24 (this box has 12 cores, each with two hyperthreads)
- Next, it creates a sample data set (to be read in the subsequent steps)
- It then uses DS2 to define a thread - like a function or a method. The thread will read a row from the sample data set, increment a count and do some arbitrary compute activity (to ensure the test exercise isn't I/O-bound, thereby preventing DS2 from showing its capabilities)
- Next, it executes DS2 again. The previously defined thread is used inside an executable piece of code (once, because threads=1). This takes 25 seconds elapsed time (and 25 seconds CPU time)
- Now we execute the same code with DS2 but we use eight threads (each processing one eighth of the input records). This takes 4.4 seconds to complete (but consuming a total of 33 seconds of CPU time across the eight threads)
- Finally, we execute a traditional DATA step to do the same task. Not surprisingly, it takes a similar amount of time to complete the task as the single-threaded DS2 code
This example suggests that great benefit can be achieved with DS2 threads. In our simple case, the code was compute-bound. If your task is more I/O-bound then the benefits may be less predictable. Jason, from the DS2 development team, recently told me:
When DS2 does I/O, it starts one "reader" thread. This reader thread requests blocks of data from a data source and puts the blocks on a queue. Then, either the DS2 "DATA" program fetches those blocks off the queue or DS2 "THREAD" programs fetch the blocks off the queue.So, as is usually the case when we talk of performance, a lot depends on your hardware architecture and the amount of effort you put into the tuning of your architecture and code. Nonethess, with DS2 it seems that there are benefits aplenty to be had.
The key here is there is one reader thread and one or more "compute" threads. As long as the reader thread can keep up with the compute threads, we should see a speed up in execution time. One reader thread works well for most data sources as data sources usually present one source to read from.
With a data source like SPDE, there could be multiple data sources across multiple devices. Right now DS2 does not take advantage of having multiple reader threads. However, I believe the architecture is flexible enough to allow mulitple reader threads.
At this point, we have been discussing I/O on a single machine. SAS High-Performance Analytics and SAS Scoring Accelerator execute DS2 in parallel across many machines. The I/O model for these types of systems is different and enables DS2 programs to better take advantage of multiple storage devices.
In my next post, I'll wrap up the topic with a few extra details.
DS2:
NOTE: DS2. Data Step Evolved?
NOTE: DS2, Learn Something New!
NOTE: DS2, SQL Within a SET Statement
NOTE: DS2, Threaded Processing
NOTE: DS2, Final Comments
NOTE: DS2, Final, Final Comments
Labels:
Performance,
SAS,
Syntax
Subscribe to:
Posts (Atom)

