Monday, April 29, 2013

httpsnow.org has a certificate problem

I'm still undecided on whether or not the Electronic Frontier Foundation are doing this to be funny.  You may know that the EFF have a project to advocate for universal use of TLS for web traffic, in support of user privacy.  They've also released a tool, HTTPS Everywhere, to provide a means for encrypted access to unencrypted websites.  One of their particularly interesting projects is the SSL Observatory, which looks at issued certificates in use on the web and evaluates them for vulnerabilities and for certificate issuance practice at certification authorities.

If you go to their TLS advocacy website, https://www.httpsnow.org, you may see something like this:



Given what the EFF is doing with HTTPS advocacy and its investigations of shoddy CA practices, I found this very surprising.  Unfortunately, however, it's common for there to be problems with web server certs, and that's the case here (i.e. it's not that there was a compromise).

What happened here is that the subject name/subject alt name is ... "*.eff.org".  So, aside from points lost for the use of a wildcard cert, the EFF are using a certificate from what's essentially, for the purpose  of certificate validation, an unrelated and incorrect certificate.  


Sunday, December 23, 2012

Not quite getting the point on trusted third party authentication

I recently learned that a former co-worker is being treated for cancer and has been communicating with friends through Caring Bridge.  So, I went to leave a note in her guest book and was given the option to log in through Facebook Connect or to create a local account.  I opted for the Facebook route (so sue me), and was taken here:


I can see where they may want to allocate local resources, but they don't seem to have quite grasped the delegated authentication thing.  It seems clear to me that there's no benefit to using Facebook to authenticate (unless you see it as a positive to have Caring Bridge post your activities to your wall [incidentally, I always set the visibility of those posts to "me, only"]).

I wonder how common it is to have app developers completely misconstrue the purpose of third-party authentication.

Sunday, May 13, 2012

Crypto tokens/smart cards

It's been awhile.

Anyway, I'm starting a project with smart cards and was quite startled to find that if they're for sale in the US it's hard to find those vendors.  I ended up ordering a couple of Feitian ePass tokens from Gooze.  Gooze is based in France and Feitian is headquartered in China.  I also ordered an Athena PKI token, and they're based in Japan (the token shipped from Japan, as well).

It's not that hard to find smart card readers in the US (to read DOD CAC cards and similar), which suggests that the problem here really is crypto export issues.  But still, why isn't someone in the US selling to American customers?

Wednesday, January 25, 2012

Indoor temperature monitoring

This is not particularly technical but might interest someone.  I live in interior Alaska and during the winter it's not that unusual for the outside temperatures to drop to -40, or lower.  Needless to say, if the indoor heating fails some really terrible things could happen to the plumbing and other household systems,  so using the software that pulls data out of my weather station and uploads it to Weather Underground I was able to throw together a trivial script to send email and an SMS message if the indoor temperature drops below 60F.

I have an Ambient Weather WS-1080, which was the predecessor to this and substantially similar.  It's pretty unsophisticated and not that accurate but it's (mostly) sufficient for my purposes and is absolutely an excellent value.  I use Tee-Boy's WeatherSnoop to pull data off the weather station console (there's a USB port) and upload it to wunderground.com.

WeatherSnoop has an option to write the data out to a sqlite database, and being curious, of course, I turned it on.  You can look at the contents of a database from the sqlite command line, and this is what I saw:

sqlite> .tables
barometricPressure  extraHumidity8      extraTemperature9   rainRate          
barometricTrend     extraHumidity9      forecast            solarRadiation    
dayRain             extraTemperature1   indoorDewPoint      uvIndex           
extraHumidity1      extraTemperature10  indoorHeatIndex     windChill         
extraHumidity10     extraTemperature2   indoorHumidity      windDirection     
extraHumidity2      extraTemperature3   indoorTemperature   windGust          
extraHumidity3      extraTemperature4   monthRain           windSpeed         
extraHumidity4      extraTemperature5   outdoorDewPoint     yearRain          
extraHumidity5      extraTemperature6   outdoorHeatIndex  
extraHumidity6      extraTemperature7   outdoorHumidity   
extraHumidity7      extraTemperature8   outdoorTemperature
sqlite>

The table "indoorTemperature" popped right out as something I could use to monitor the house when I was away from home, so I took a look at the schema and found that there are two values: time and value:

sqlite> .schema indoorTemperature
CREATE TABLE indoorTemperature ( 'time' INTEGER, 'value' FLOAT );
sqlite>

So, I just wrote a script to dump the indoor Temperature table
#!/bin/ksh

WEATHERFILE=/Users/melinda/Documents/weather.db

sqlite3 $WEATHERFILE <<EOF 
.separator " "
select strftime('%Y:%m:%d:%H:%M', time, 'unixepoch', 'localtime'),value from indoorTemperature;
EOF

What this does is use a Unix "here document" to provide scripted commands to the sqlite command line.  .separator " " changes the character used to separate fields on output to a space (by default it's a vertical bar, or pipe character).

The select statement is actually pretty straightforward and kind of inefficient and brainless, to be honest.  Basically I'm selecting every record and formatting the time for printing.

This little routine is invoked by the real script, which looks like this:

#!/bin/ksh

CURTMP=`/Users/melinda/lib/curindoortemp | tail -1 | cut -d" " -f2`
ADDR=0000000000@msg.acsalaska.com,xxxxxxx@gmail.com
THRESHHOLD=60

if [ $CURTMP -lt $THRESHHOLD ]
then
        Mail $ADDR <<EOF
        Temperature is $CURTMP
EOF
fi

The curindoortemp script is the one described above.  I grab the last record and print the second field, which is the temperature.  If it's less than the thresshold value (in this case, 60 degrees Fahrenheit) I send off email to the addresses listed in the ADDR field.  Obviously it would be far better programming practice to have the curindoortemp script just return the temperature from the last record.

Wednesday, April 20, 2011

On the perils of not knowing some history ...

I followed a tweeted link on the use of OpenBSD in a corporate environment, and listed on that page as a "Featured article" was this piece on OpenSSH configuration.  It's almost entirely reasonable stuff, but the author added a note that "Bob" had an excellent point:
Saying "don't login as root" is h******t. It stems from the days when people sniffed the first packets of sessions so logging in as yourself and su-ing decreased the chance an attacker would see the root pw, and decreast the chance you got spoofed as to your telnet host target, You'd get your password spoofed but not root's pw. Gimme a break. this is 2005 - We have ssh, used properly it's secure. used improperly none of this 1989 will make a damn bit of difference. -Bob
That's actually a terrible point, and based at least in part with a lack of familiarity with both history and best practices.  It's not because of fear of sniffers, and I really don't think it ever has been, particularly.  Here's why you prevent root logins over the network:
  1. It increases (or should) the amount of entropy in the system.  If you allow root logins from the network an attacker has to guess one password.  If you allow users who can su or sudo in from the network, the attacker has to guess who they are (not that difficult if it's a familiar system) and then guess/break the user's password and the root password.  In the meantime it's not that much extra effort to log in as yourself and then su.  That's a pretty good security/UX tradeoff, as these things go.
  2. We humans have been known to make mistakes a time or two and a fat-finger error when you're root is potentially far more consequential than the same mistake made when you don't have elevated privileges.  In the interest of getting work done it's sometimes more efficient to run a root shell, but in general, best practices says don't use elevated privileges when you don't need them.
Come to think of it, "Bob's" comment may stem from the well-known tendency to confuse security with encryption.  Good credentials and good security practices are critical, too.

Friday, June 12, 2009

Kernel dives vs. sysctl

A few weeks ago I decided to take a look at porting a multiprocessor xcpustate to FreeBSD 7.2[*]. It compiled but didn't actually run and it failed less than gracefully. It turned out that xcpustate was doing a kmem dive to pull utilization stats out of the kernel but that it wasn't detecting that the variable it was trying to find is no longer there. Let me back up a little.

Back in the day, we used to collect system performance data, in most cases, by grotting around in kernel memory. It worked more-or-less like this:
  1. identify the kernel variable containing the data you want
  2. do an nlist(3) to get the address of the data
  3. open /dev/kmem
  4. lseek(2) to the address returned by nlist
  5. read(2) the data
This assumed that the person writing the utility had some knowledge of the kernel and kernel data structures. To my knowledge the first operating system that supported some sort of direct copyout() of some kernel variables was Unicos (Cray's version of Unix for their supercomputers; I used it on the X/MP version). It eventually became ubiquitous as sysctl(3). To do the same job (grabbing the contents of a kernel variable) became as simple as this:
  1. identify the sysctl-supported variable you want
  2. call sysctlbyname(3)
Here's the big surprise: the conventional wisdom has always been that lseek()ing around kmem was costly and that it would be cheaper to just copy the data out directly (although copyout() isn't free). Out of curiosity, while waiting for a phone call today I found a variable that both was accessible through sysctl and appeared in the kernel symbol table: maxfilesperproc and implemented grabbing it in both styles. What I found surprised me a bit, which was that the nlist/lseek/read version consistently took about half the time of a sysctlbyname. The sysctl man page does warn that converting the string form of the name to the mib form can be costly and I'll assume for now that that's where the time is being lost. I'll experiment with constructing the mib by hand later.

Update: I reinstrumented my code and found a bug (oh no!) in the stuff to format printing microsecond timings, and what I found was: 1) the version that just does a sysctlbyname() runs two orders of magnitude (that's a lot!) faster than the code that does a kernel dive, and 2) the code that uses a preformatted/loaded mib and calls sysctl instead of sysctlbyname runs about 4 times faster than using sysctlbyname, or about 400 times faster than doing a kernel dive. That's a little dramatic.

Also, note that the memory scrape approach can be extended to read the contents of variables in executing programs other than the kernel. Start with reading the proc table from /dev/kmem to get the address space of the process of interest, get the name list out the executable (if you've ever wondered why strip(1) exists ... ) and go forward from there.

[*]All of this, of course, doesn't address the question of whether or not FreeBSD is instrumented to be able to pull per-processor performance data out of the kernel, which it appears not to be.

Monday, June 1, 2009

Facebook performance

I continue to be surprised by Facebook's performance problems. Not just that they're not improving, but that they're actually getting worse. Maybe that's to be expected when you've got a highly interactive system with millions of active users that's written in a scripting language (PHP). But just saying "It's written in PHP" isn't really enough - it could be that the PHP interpreter has performance problems, or the scripts themselves could be inefficient.

And that doesn't even begin to address what are essentially data center questions. There may be insufficient computing or network resources, what resources they have may be badly allocated or administered, etc. I once worked on a hierarchical storage management system and there were users complaining that it was taking as long as 1/2 hour to open a file. I turned on disk tracing (this was AIX on an RS6k) and found that namei() on files in one particular directory were causing a single disk head to skip all over the place. From there it was easy to determine that the directory in question contained an insane number of files and there were indirect indirect indirect indirect ... blocks to hold the directory. By simply reorganizing the filesystem - striping it across multiple platters and multiple spindles and relieving that one poor disk head from having to do all the work - I was able to speed up lookups by three orders of magnitude. That is to say, I was able to get lookups down to 1/1000 of what they had been taking without changing a single line of code or forcing the user to change their (unfortunate) behavior. I suspect that there are lots of instances of those sorts of problems lurking inside the Facebook data center and I'm not sure that the operating systems they're using are instrumented well enough to be able to isolate them.

But still, I don't think sending over a megabyte of scripts down to the user's browser has nothing to do with it. What a mess.