Squashed commit of the following:
commit 8fcdd34f87fe9daaac235315d3b6e1d4986ea0ed
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Nov 5 14:26:57 2013 +0100
Fix for ramdisk generation (?)
Change-Id: If7fd55d9a146d7d7b621da274ac0472a5b246c36
commit 5ffd24ad49bff556adceb050c7cbfddd774ed200
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Nov 5 13:16:56 2013 +0100
Rebase fixes
Too lazy to merge this into the appropriate patch.
Change-Id: I66505d6a0d65330656e78ec9615a026f892b8609
commit 0240220770e6cc18fe16bd0c82e6e5e2c88a4301
Author: David van Moolenbroek <david@minix3.org>
Date: Sun Oct 6 15:58:54 2013 +0200
VFS: further cleanup of device code
- all TTY-related exceptions have now been merged into the regular
code paths, allowing non-TTY drivers to expose TTY-like devices;
- as part of this, CTTY_MAJOR is now fully managed by VFS instead of
being an ugly stepchild of the TTY driver;
- device styles have become completely obsolete, support for them has
been removed throughout the system; same for device flags, which had
already become useless a while ago;
- device map open/close and I/O function pointers have lost their use,
thus finally making the VFS device code actually readable;
- the device-unrelated pm_setsid has been moved to misc.c;
- some other small cleanup-related changes.
Change-Id: If90b10d1818e98a12139da3e94a15d250c9933da
commit 07d10dc91955475340005fa86adf71aa981712c8
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 14:59:59 2013 +0200
IS: dump number of in-use FDs for VFS
Change-Id: If0e2092d5a8c384c31b1f44cc0591bb119c6d8de
commit ced39723e3e7c39a4a2ef516834702a36a6a26c8
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 01:40:58 2013 +0200
TTY: fix earlier PTY select "improvement"
It was just plain wrong.
Change-Id: Ieab4b4f01d9461e05e0d0ba6427a99d863d6b98d
commit a4a63c15cc5f9daf2e149b84ec930834fe24a983
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 16:31:35 2013 +0200
Extend dupfrom(2) into copyfd(2)
This single function allows copying file descriptors from and to
processes, and closing a previously copied remote file descriptor.
This function replaces the five FD-related UDS backcalls. While it
limits the total number of in-flight file descriptors to OPEN_MAX,
this change greatly improves crash recovery support of UDS, since all
in-flight file descriptors will be closed instead of keeping them
open indefinitely (causing VFS to crash on system shutdown). With the
new copyfd call, UDS becomes simpler, and the concept of filps is no
longer exposed outside of VFS.
This patch also moves the checkperms(2) stub into libminlib, thus
fully abstracting away message details of VFS communication from UDS.
Change-Id: Idd32ad390a566143c8ef66955e5ae2c221cff966
commit 663e0abb01990da768f1701df6366a4d3bbd67c1
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 12:50:44 2013 +0200
VFS: better dupfrom(2) deadlock detection
Change-Id: I29f00075698888c7c8ca60b47ab82fba8c606f4e
commit 54c509663718cc109b0f11ab0adc196bfa82015b
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Oct 4 21:10:01 2013 +0200
UDS: sendmsg/recvmsg fixes
- sendmsg: the accumulation of multiple in-flight file descriptors was
already described in the comments; now the code actually does what
the comments say :) -- also, added robustness in case of a failure;
- recvmsg: only create a socket rights message if there are file
descriptors pending at all;
- recvmsg: copy back the control message length;
- recvmsg: use CMSG_SPACE instead of CMSG_LEN to compute sizes.
Not sure if all of this is now working according to specification,
but at least tmux seems to be happy with it.
Change-Id: I8d076c14c3ff3220b7fea730e0f08f4b4254ede5
commit b0a891c329d721eb84e06ecac803789fe8197306
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Oct 4 18:41:21 2013 +0200
UDS: add support for FIONREAD
Change-Id: I50030012b408242a86f8c55017429acdadff49d1
commit e81d59c874d57d0d0010d3b1ace2351d01f66766
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Oct 4 18:00:42 2013 +0200
UDS: align struct sockaddr_un with NetBSD
Well, make a start, anyway. Our copy was missing a legacy field from
the structure, that could very well cause applications to fail trying
to set, clear, or check it. As a consequence, SUN_LEN now yields the
same result as on NetBSD.
Change-Id: I80f6aff7769be402b3bd3959f64d314509ed138c
commit 66f8ad9ac7f0e008cdf416877eeab63ed599477d
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Oct 4 17:57:42 2013 +0200
UDS: support for nonblocking sockets
This patch includes several other fixes, which are now tested in the
test56 test set.
Change-Id: I9535d5a6c072abf966252838522c5f65b353c6c2
commit 9b4a7ba7ced0b40bf650451ec7baf2935b2ccfe7
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Oct 4 16:46:18 2013 +0200
UDS: clean up source code
- move VFS calls to a separate source file;
- solve a few subtle bugs, mostly in error handling;
- simplify debug reporting code;
- make a few definitions more independent;
- restyle to something closer to KNF.
Change-Id: I7b0537adfccac8b92b5cc3e78dac9f5ce3c79f03
commit 881641da1fa711d2d4bf8687cb5beb843801daf8
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Oct 4 16:29:40 2013 +0200
UDS: split off from PFS
Change-Id: I769cbd64aa6e5e85a797caf0f8bbb4c20e145263
commit 2baef79e569c4cebda1a3f831b8e329295e449e6
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Oct 2 19:37:00 2013 +0200
at_wini: PCI-only now; one controller per instance
- remove non-PCI support, since all supported platforms with at_wini
devices also have PCI support by now;
- correspondingly, stop using information from the BIOS altogether;
- limit each driver instance to one controller, to be in line with
the general MINIX3 one-instance-per-controller driver model; this
limits the number of disks per at_wini instance to four;
- go through the controllers by the order of their occurrence in the
PCI table, thus removing the exception for compatibility devices;
- let the second at_wini instance shut down silently if there is only
one IDE controller;
- clean up some extra code we don't need anymore, and resolve some
WARNS=5 level warnings.
Overall, these changes should simplify automatic loading of the right
disk drivers at boot time in the future.
Change-Id: Ia64d08cfbeb9916abd68c9c2941baeb87d02a806
commit f7aaabe7e0ae93b439b02d3c21c38f11bae7bbcf
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Oct 1 00:42:41 2013 +0200
system.conf: subsystem VID/DID matching support
- change "vid/did" to "vid:did", old form still supported for now;
- allow "vid:did/subvid:subdid" specification in system.conf, in
which case a device will be visible to a driver if the subsystem
VID/DID also match.
Change-Id: I7aef54da1b0bc81e24b5d98f1a28416f38f8b266
commit 178d896a32140285e7e9f00dba7bb4a871752f60
Author: Lionel Sambuc <lionel@minix3.org>
Date: Tue Oct 8 19:00:16 2013 +0200
usr.bin/man: Update
Change-Id: I0c5d2115ba384687032f7b2af50d99dedc323b7a
commit 33e8e00217bbf38300f266dda5a494ad4e51bebd
Author: Lionel Sambuc <lionel@minix3.org>
Date: Tue Oct 8 18:05:29 2013 +0200
external/bsd/mdocml: Update
Change-Id: I17b54e52e8322676d83ed4386f586f8ef3029f72
commit e28c0794f18eae77faf3055dadbe76a963242c4e
Author: Lionel Sambuc <lionel@minix3.org>
Date: Thu Oct 3 18:26:21 2013 +0200
Adapting build system to call MAKEDEV for /dev
* Remove static proto.dev
* Update releasetools/*image.sh not to use proto.dev, as well as
minor comments cleanup
* Add TOOL_TOPROTO
Change-Id: If7dc16d4ebb3b0c4e859786fad25d4af000c999f
commit d26b8149f47461aac67078b8a63ddae21ecdf4f8
Author: Lionel Sambuc <lionel@minix3.org>
Date: Thu Oct 3 13:54:24 2013 +0200
MAKEDEV: Add mtree output, and ramdisk set.
Change-Id: I36cb7e9451960189a33a04a5c2e3ddb19c7be75e
commit 8dc3249c2ec209689922abb306b0e3d15e6a0347
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 20:24:25 2013 +0200
3.3.0 sync fixes
Keeping this version at 3.2.1 though, mainly for pkgsrc.
Change-Id: Ic91e326623f6dbad58754152c4a632e4406f1117
commit f61a525630a33734481bfd8f0994ae33fef6c3fb
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Oct 2 13:37:49 2013 +0200
keymaps: improve keypad slash support
Now that the keymaps can distinguish between the regular slash key
and the slash key on the numeric keypad, we can avoid localization
of the latter.
Change-Id: I20ead7d26a9baa82f5a522562524fd75d44efb42
commit 81e1e0e8e10c08459af9b7d24f51eff41225b62d
Author: Mikal Villa <mikal.villa@gmail.com>
Date: Wed Oct 2 12:30:03 2013 +0200
Norwegian keymap
Change-Id: I181234afc8f1a058e92af6c1fe88979463aaff45
commit c9fc72ef1769f677ad6c932e84959f8756b32264
Author: Thomas Cort <tcort@minix3.org>
Date: Fri Sep 6 21:40:42 2013 -0400
Importing usr.bin/uname
Change-Id: I4c316221e288edd839e26a2af4cb59f28bf722c1
commit 320e18fd3269219d4b68ba0df3ff8f8434a24a09
Author: Thomas Cort <tcort@minix3.org>
Date: Fri Sep 6 21:40:31 2013 -0400
uname: normalize release and version
Most systems provide the full version number in the
'release' field and the kernel version in 'version'.
Minix used to split the full version number between
release and version which caused problems for pkgsrc
and other applications. This patch brings Minix's
uname in line with other systems such as NetBSD.
It also brings the getty banner in line with NetBSD.
Old Minix uname:
sysname->Minix
nodename->10.0.2.15
release->3
version->2.1
machine->i686
New Minix uname:
sysname->Minix
nodename->10.0.2.15
release->3.2.1
version->Minix 3.2.1 (GENERIC)
machine->i686
Change-Id: I966633dfdcf2f9485966bb0d0d042afc45bbeb7d
commit 3840323699856b0ab67b3c6c3b79695acc25ffcb
Author: Lionel Sambuc <lionel@minix3.org>
Date: Thu Sep 19 10:57:10 2013 +0200
Replacing timer_t by netbsd's timer_t
* Renamed struct timer to struct minix_timer
* Renamed timer_t to minix_timer_t
* Ensured all the code uses the minix_timer_t typedef
* Removed ifdef around _BSD_TIMER_T
* Removed include/timers.h and merged it into include/minix/timers.h
* Resolved prototype conflict by renaming kernel's (re)set_timer
to (re)set_kernel_timer.
Change-Id: I56f0f30dfed96e1a0575d92492294cf9a06468a5
commit 23906ce75094c89071c04c6a2a2da780d5d1d9d1
Author: Lionel Sambuc <lionel@minix3.org>
Date: Wed Oct 2 10:56:24 2013 +0200
ARM serial driver: Comment termios_baud_rate.
The B0-B115200 defines are flags, and not the actual speed they
represent.
This fixes an incoherency for B0 handling, and documents why it is
required to call the function again after changing the speed flag.
DFL_BAUD is set to one of the flag, so to translate it to an actual
speed, the function calls itself again, which will always be able to
finish without inducing another recursive call.
Change-Id: I04ebfaefee31a88d05f0b726352d1581a966147b
commit 92db382b054926bc494cf369731e67af91033cc8
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Oct 1 23:25:01 2013 +0200
TTY: skip /dev/log checks if console is serial
It is unclear why /dev/log has its own open/close rules, but those
rules conflict with serial console redirection. This does not solve
the root of the problem, but it puts back in place more or less the
same workaround that was already in place before the TTY overhaul.
Change-Id: Ib53abbc28a76c1f2b0befc8448aeed0173bc96a5
commit f1b6f94f744e82b0a4f0fbe8967a4986e079c7cd
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 19:45:42 2013 +0200
Corrections to keep VFS ABI-compatible
Change-Id: Id8a53a687577c1b6f83c11c2e5fb0b8ed1cd2a08
commit 1d9f12eec8edb0ed8b9bc3cb05eb342d6e4ffb03
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 19:35:11 2013 +0200
Remove support for obsolete 3.2.1 ABI
Change-Id: I76b4960bda41f55d9c42f8c99c5beae3424ca851
commit 93267010fcf5a6de5bc4ea511e1edab09c2362aa
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Aug 31 22:59:44 2013 +0200
Fix various file system warnings
Change-Id: Ied10498c3ae14f9f2fd06914f23239df330fa296
commit 42a4fca69efd41f11d28e5b4a4df45c3a14ea91f
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 19:22:13 2013 +0200
VFS/FS: replace protocol version with flag field
The main motivation for this change is that only Loris supports
multithreading, and Loris supports dynamic thread allocation, so the
number of supported threads can be implemented as a bit flag (i.e.,
either 1 or "at least as many as VFS has"). The ABI break obviates the
need to support file system versioning at this time, and several
other aspects are better implemented as flags as well. Other changes:
- replace peek/bpeek test upon mount with FS flag as well;
- mark libsffs as 64-bit file size capable;
- remove old (3.2.1) getdents support.
Change-Id: I313eace9c50ed816656c31cd47d969033d952a03
commit 6237d4580b2e7f4ca37222581c44f3898a9f2985
Author: Lionel Sambuc <lionel@minix3.org>
Date: Fri Apr 19 13:16:51 2013 +0200
usr.bin/stat Update
Change-Id: I029160c73baab1b3465bc5397a36c55886db225b
commit 9e7f974e9fa61e0e1d9cb519a69155eb71826f9d
Author: Lionel Sambuc <lionel@minix3.org>
Date: Fri Apr 19 09:54:51 2013 +0200
almost aligned ioctl prototype
Change-Id: I7f3eaa99d2a9767f71e8387cea5c7f56dcb28f99
commit b1a9440566584280bd07c6e646bb18d5b8133786
Author: Lionel Sambuc <lionel@minix3.org>
Date: Thu Apr 18 14:42:15 2013 +0200
32 to 64 bits fsblkcnt_t and fsfilcnt_t.
Change-Id: I432229143c85cd178262b802a76ac606801ac59a
commit 8dea53df7843cf4a96601e019f941e7d13a878c0
Author: Lionel Sambuc <lionel@minix3.org>
Date: Thu Apr 18 11:08:16 2013 +0200
moving prototypes to lib.h
Change-Id: If53d3f5ee761b10e0f3d4346a0c5b39ba7901c65
commit 0d00df150181bcc7ee20726801bbc34ebe978a29
Author: Lionel Sambuc <lionel@minix3.org>
Date: Thu Apr 18 11:07:44 2013 +0200
struct uucred
Change-Id: Ia97cb6c38bb566be30d568a252ae7b76142a21dd
commit 1e355e9d217877943e01c5b8859d66258e54b3e5
Author: Lionel Sambuc <lionel@minix3.org>
Date: Tue Apr 16 11:48:54 2013 +0200
Alignement on netbsd types, part 1
The following types are modified (old -> new):
* _BSD_USECONDS_T_ int -> unsigned int
* __socklen_t __int32_t -> __uint32_t
* blksize_t uint32_t -> int32_t
* rlim_t uint32_t -> uint64_t
On ARM:
* _BSD_CLOCK_T_ int -> unsigned int
On Intel:
* _BSD_CLOCK_T_ int -> unsigned long
bin/cat is also updated in order to fix warnings.
_BSD_TIMER_T_ has still to be aligned.
Change-Id: I2b4fda024125a19901120546c4e22e443ba5e9d7
commit e7b39757eac1ab7aac5a59d7aa124fc316a6e01f
Author: Lionel Sambuc <lionel@minix3.org>
Date: Fri Aug 23 20:27:27 2013 +0200
Adapt the type used for adjtime_delta
clock_t is currently a signed type, but in NetBSD this is not the
case. As we plan on aligning our types we have to change this as this
prevents negative delta from being correctly used.
Change-Id: I9bccdee2b41626b0262471dc1900de505a1991a7
commit 5a551dbe663b22760632c5232326e4101360b330
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 19:01:19 2013 +0200
Butchered version of "Bumping version.."
Change-Id: Ieb5f31fe037bcf45cc86a906f983a0ab3f636468
commit 6d601ae1d7072a8fb6cd908ff1d24ea4ca9bf1b0
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 18:57:04 2013 +0200
Trimmed version of "VFS: use 64-bit file..."
Change-Id: I1b53ef05e939f6afed97f507d443a1d60ed81a21
commit 7adaa6556051b3800670399ee26d246fa9f32486
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Oct 5 18:37:39 2013 +0200
Trimmed version of "stat.h: remove some big_ types"
Change-Id: I6c17d6e5e9d2ab4c1bbf33e461cec4c37b870266
commit 3da9bd17116b66416e8e10ce148599b8ea723680
Author: Thomas Veerman <thomas@minix3.org>
Date: Thu Mar 7 14:46:21 2013 +0000
Define protocol version of {mode,ino,uid,gid}_t
Change-Id: Ia2027749f2ce55a561d19eb895a5618505e9a2ac
commit 6ca17cd644c3b3e4f318faf7d1242855af473b6f
Author: Thomas Veerman <thomas@minix3.org>
Date: Mon Mar 25 21:09:10 2013 +0000
VFS-FS protocol: add versioning
Change-Id: Ice6fbfd4b535b7435653fa08b27a3378d1cfdbf8
commit a7ff601df7d61c0925241ce10aee3a7e99343b3a
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Sep 28 14:46:21 2013 +0200
Input infrastructure, INPUT server, PCKBD driver
This commit separates the low-level keyboard driver from TTY, putting
it in a separate driver (PCKBD). The commit also separates management
of raw input devices from TTY, and puts it in a separate server
(INPUT). All keyboard and mouse input from hardware is sent by drivers
to the INPUT server, which either sends it to a process that has
opened a raw input device, or otherwise forwards it to TTY for
standard processing.
Design by Dirk Vogt. Prototype by Uli Kastlunger.
Additional changes made to the prototype:
- the event communication is now based on USB HID codes; all input
drivers have to use USB codes to describe events;
- all TTY keymaps have been converted to USB format, with the effect
that a single keymap covers all keys; there is no (static) escaped
keymap anymore;
- further keymap tweaks now allow remapping of literally all keys;
- input device renumbering and protocol rewrite;
- INPUT server rewrite, with added support for cancel and select;
- PCKBD reimplementation, including PC/AT-to-USB translation;
- support for manipulating keyboard LEDs has been added;
- keyboard and mouse multiplexer devices have been added to INPUT,
primarily so that an X server need only open two devices;
- a new "libinputdriver" library abstracts away protocol details from
input drivers, and should be used by all future input drivers;
- both INPUT and PCKBD can be restarted;
- TTY is now scheduled by KERNEL, so that it won't be punished for
running a lot; without this, simply running "yes" on the console
kills the system;
- the KIOCBELL IOCTL has been moved to /dev/console;
- support for the SCANCODES termios setting has been removed;
- obsolete keymap compression has been removed;
- the obsolete Olivetti M24 keymap has been removed.
Change-Id: I3a672fb8c4fd566734e4b46d3994b4b7fc96d578
commit 0a3e629c0d4101a33c8b2e79c2e52d312978d34b
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Sep 27 11:56:29 2013 +0000
TTY: allow selecting on translated minors
Due to the existence of /dev/console and /dev/log, and the new
"console=" setting, it is now possible that a single non-PTY object
(e.g. serial) is accessible through two different minor numbers. This
poses a problem when sending late select replies (CDEV_SEL2_REPLY),
because the object's minor number can not be used to identify the
device. Since selecting on such objects through translated minor
numbers is actually required, we now save the minor number used to
initiate the select query in order to send a late reply.
The solution is suboptimal, as it is not possible to use two different
minors to select on the same object at once. In the future, there
should be at least one select record for each minor that can be used
with each object.
Change-Id: I4d39681d2ffd68b4047daf933d45b7bafe3c885e
commit fe2167a148dd17e15502d98e64057239fa122dd5
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Sep 21 17:35:15 2013 +0200
Take LOG out of the boot image
Change-Id: Id2629776b53aae46629b04a42c15cbbacac9b949
commit 81e138bb5ba168322eb780978ed4a90fc1ab7945
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Sep 21 15:03:20 2013 +0200
Kernel: make SIGKMESS target process list dynamic
The set of processes to which a SIGKMESS signal is sent whenever new
diagnostics messages are added to the kernel's message buffer, is now
no longer hardcoded. Instead, processes can (un)register themselves
to receive such notifications, by means of sys_diagctl().
Change-Id: I9d6ac006a5d9bbfad2757587a068fc1ec3cc083e
commit 9afa18ec208646a9244048a1d444c474ceb47c73
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Sep 21 00:58:46 2013 +0200
Rename SYSCTL kernel call to DIAGCTL
Change-Id: I1b17373f01808d887dcbeab493838946fbef4ef6
commit f82f1ec4b06240e96ed8e64e4bfdf5b06d688f89
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 18 20:13:59 2013 +0200
Prevent Jenkins from breaking on old df(1)
Change-Id: Ie8e1758bd9de4d5c95c597302f80d568058fbb68
commit 8a430b673dc97767306ddc9cb1a5367f333b7602
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 18 14:09:47 2013 +0200
Add testvnd.sh test script
As part of this, change the "run" script to allow certain scripts to
be run as root only.
Change-Id: I846e41037f9d4f6c7fc0b5ea8250303a7bd72f5d
commit fd57902664f398edd9274c2ab1d87daa0b4f4584
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 18 14:06:46 2013 +0200
Import NetBSD vndconfig(8)
The tool has been changed heavily to match our VND driver model.
NetBSD is in the process of renaming it from vnconfig(8) to
vndconfig(8). To keep things in sync, we have to play along.
Change-Id: Ie86df184f03ab00573ea76b43c9caa0412e8321d
commit 6809baec22122124a04d8bf8f4f8b5fc7778cb3d
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 18 14:02:17 2013 +0200
Add VND driver, providing loopback devices
Change-Id: I40fa695e28c67477a75383e6f1550e451afcab41
commit 4cb7efbfd17b934f914f633c8bebf5bec7934e5d
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 18 13:55:15 2013 +0200
VFS: add dupfrom(2) call
This call copies a file descriptor from a remote process into the
calling process. The call is for the VND driver only, and in the
future, ACLs will prevent any other process from using this call.
Change-Id: Ib16fdd1f1a12cb38a70d7e441dad91bc86898f6d
commit 378f893b1e710563e2ef2eee457255a6f3037885
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 18 11:46:08 2013 +0200
tests: do not skip installed shell tests
When installed, the test scripts lose their ".sh" suffix, causing them
to be skipped by the "run" script. With this patch, the tests are no
longer specified with ".sh" suffix in the run script, and the suffix
is added automatically as necessary.
Change-Id: I0b72312e79992b9818559c6546a0e52cd95184c2
commit ecdd53c03077c3cceb29485aa56021d99485d6a5
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 18 13:41:56 2013 +0200
blocktest: prepare to be run as part of tests
- fail SEF initialization if any of the subtests failed, so that the
party invoking the "service up" can tell whether the test succeeded;
- add "nocontig" option, because VM isn't particularly good at
allocating contiguous memory;
- add "silent" option, because it floods the console otherwise;
- allow the device size to be smaller than the maximum transfer size;
- install files to installed test directory.
Change-Id: I45c818f817c11d90c5f94ae26a2fc49e36e6761e
commit 848ef711847b7adcf20dce333435802c4e1018ea
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 18 13:38:36 2013 +0200
Enable devname(3)
There is no support for a device name database yet, so this call is
expected to be fairly slow.
Change-Id: I73aa5f267e2b6921b7d3bbdcc4beac463931132c
commit b84199849ad72030993bcf3f943da1509be15ded
Author: David van Moolenbroek <david@minix3.org>
Date: Sun Sep 15 13:09:00 2013 +0200
libbdev: be less noisy about clean driver restarts
Change-Id: Ie02a459c9b544d361ab00bac431ef99de53b0c5f
commit 4f01352a2a6c82fc11465bc555c25e0344c5b970
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Sep 14 14:43:53 2013 +0200
Straighten ioctl.h
- include all ioctl subheaders, properly listing all letters;
- change FBD's ioctl calls to use 'B' instead of 'F', in
preparation of the VND driver.
Change-Id: Ia718979568cc057f47cf505a89238d5b3b6695d4
commit 0ee3f3aeb3d634133a9cd8c3b81589772235c617
Author: David van Moolenbroek <david@minix3.org>
Date: Sun Sep 15 18:55:42 2013 +0200
VM: readd support for forgetting cached FS blocks
Not all services involved in block I/O go through VM to access the
blocks they need. As a result, the blocks in VM may become stale,
possibly causing corruption when the stale copy is restored by a
service that does go through VM later on. This patch restores support
for forgetting cached blocks that belong to a particular device, and
makes the relevant file systems use this functionality 1) when
requested by VFS through REQ_FLUSH, and 2) upon unmount.
Change-Id: I0758c5ed8fe4b5ba81d432595d2113175776aff8
commit c1652705b625a978ea2af169a21349352bd6575f
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 11 14:50:18 2013 +0200
filter: use libblockdriver
Change-Id: Ifbca2482e996ddca58036d45f557165e636fb3fa
commit ad99ccdb45b897afa7573f75b1a29c4a19767870
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Sep 10 20:25:01 2013 +0200
Rewrite character driver protocol
As a side effect, remove the clone style, as the normal device style
supports device cloning now.
Change-Id: Ie82d1ef0385514a04a8faa139129a617895780b5
commit 4f91091ea26022877fd8c29094008d8072cf5e2b
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Sep 10 16:06:37 2013 +0200
Remove support for reopening character devices
Previously, VFS would reopen a character device after a driver crash
if the associated file descriptor was opened with the O_REOPEN flag.
This patch removes support for this feature. The code was complex,
full of uncovered corner cases, and hard to test. Moreover, it did not
actually hide the crash from user applications: they would get an
error code to indicate that something went wrong, and have to decide
based on the nature of the underlying device how to continue.
- remove support for O_REOPEN, and make playwave(1) reopen its device;
- remove support for the DEV_REOPEN protocol message;
- remove all code in VFS related to reopening character devices;
- no longer change VFS filp reference count and FD bitmap upon filp
invalidation; instead, make get_filp* fail all calls on invalidated
FDs except when obtained with the locktype VNODE_OPCL which is used
by close_fd only;
- remove the VFS fproc file descriptor bitmap entirely, returning to
the situation that a FD is in use if its slot points to a filp; use
FILP_CLOSED as single means of marking a filp as invalidated.
Change-Id: I34f6bc69a036b3a8fc667c1f80435ff3af56558f
commit 8dfbb8fbe4f37642ac9700cc46ad0588687aa3a4
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Sep 10 12:19:08 2013 +0200
VFS: rework device code
- block the calling thread on character device close;
- fully separate block and character open/close routines;
- reuse generic open/close code for the cloning case;
- zero all messages to drivers before filling them;
- use appropriate types for major/minor device numbers.
Change-Id: Ia90e6fe5688f212f835c5ee1bfca831cb249cf51
commit 9fabdb8b4edf637b47f1530bea145d68a028c319
Author: David van Moolenbroek <david@minix3.org>
Date: Mon Sep 9 00:04:12 2013 +0200
Make PFS backcalls regular VFS calls
- prefix them with VFS_ as they are going to VFS;
- give these calls normal call numbers;
- give them their own set of message field aliases;
- also make do_mapdriver a regular call.
Change-Id: I2140439f288b06d699a1f65438bd8306509b259e
commit 8fbd5d36fa175e2c31ed4a239be3f785e2995aef
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 11 12:48:10 2013 +0200
TTY: use libchardriver; clean up
- writing to a PTY master side blocks if there is not already a
blocked reader on the slave side, and select now reflects this;
- internally, TTY now uses a test based on "caller != NONE" rather
than "grant != GRANT_INVALID" to identify whether a call is
currently ongoing;
- "offset" fields have been removed as they equal the corresponding
"cum" fields;
- improved variable typing and function naming here and there;
- various other small fixes.
Change-Id: I6b51452888942e864b4e034e8c8490576184a23e
commit 53df30e2b876c03b26f8f52abeafcf8f601e6fd2
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 11 01:13:59 2013 +0200
VFS: select(2) fixes
- check each file descriptor's open access mode (filp_mode);
- treat an error returned by a character driver as a select error;
- check all filps in each set before finishing select;
- do not copy back file descriptor sets if an error occurred;
- remove the hardcoded list of supported character major devices,
since all drivers should now be capable of responding properly;
- add tests to test40 and fix its error count aggregation.
Change-Id: I57ef58d3afb82640fc50b59c859ee4b25f02db17
commit f08c75810f26adf75a757b5b65fdfa2589cf701d
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 11 01:07:28 2013 +0200
Retire EBADIOCTL in favor of ENOTTY
Change-Id: I6bd0e301d21ab7f2336e350e7e6e15d238c2c93d
commit 91b7de8257c92283648c796d882a9095a1f64707
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 4 15:42:36 2013 +0000
libnetsock: use libchardriver
Change-Id: Ia5b780cad0b0c636db9bd866c7223da0d38ef6ea
commit 6a3a64d9669c5b1d6ca9acc12f02cd2f21f22d0f
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 11 00:50:36 2013 +0200
LWIP: move chardev message parsing into libnetsock
Change-Id: Ie23fd003c9fa35811548f388c8e9b55e8d9de8d7
commit 39aa2794e04ad34aecc6dcf607c0987dad13c1ef
Author: David van Moolenbroek <david@minix3.org>
Date: Mon Sep 9 14:53:28 2013 +0000
INET: use libchardriver
Change-Id: Icf8a1a5769ce0aede1cc28da99b9daf7d328182c
commit 7c0d3f508d8e840a2d30da4d3e3af696b9ad2abc
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Sep 3 02:00:20 2013 +0200
PFS: use libchardriver; clean up
- simplify and repair UDS request handling state machine;
- simplify interface used between internal modules;
- implement missing support for nonblocking I/O;
- fix select implementation;
- clean up global variables.
Change-Id: Ia82c5c6f05cc3f0a498efc9a26de14b1cde6eace
commit 5598aedb79b34fd02b3b81427d683014a3303e0c
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Sep 3 01:59:20 2013 +0200
libaudiodriver: use libchardriver
Change-Id: I299d58d110ad14b69076276ba46c4325875c34ca
commit 2aadc0aa253427d18c1e3db597c33dd06d9c380e
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Sep 3 01:57:41 2013 +0200
printer: use libchardriver
Change-Id: Ifa3cabeada74c32df2613b8279c00f86e831c775
commit aeaa28b7c68aced329d9cf5b4a01af07b2f3638e
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Sep 3 01:49:38 2013 +0200
libchardriver: full API rewrite
The new API now covers the entire character driver protocol, while
hiding all the message details. It should therefore be used by all
new character drivers. All existing drivers that already made use of
libchardriver have been changed to use the new API.
As one of the most important API changes, support for scatter and
gather transfers has been removed, as several key drivers already
did not support this, and it could be supported at the safecopy
level instead (for a future readv/writev).
Additional changes include:
- respond to block device open requests to avoid hanging VFS threads;
- add support for sef_cancel.
Change-Id: I1bab6c1cb66916c71b87aeb1db54a9bdf171fe6b
commit 6fbd39f8a65178d8fafa260dc42fde5341eb9097
Author: David van Moolenbroek <david@minix3.org>
Date: Thu Aug 1 18:20:33 2013 +0200
blocktest: add support for no alignment
Some block drivers do not impose any alignment requirements, and this
patch allows such block drivers to pass the test set. As a side effect,
minimal support for min_write is added, but this part of blocktest is
in need of further improvement.
commit d70dc11950991506fa08d144067de3fe86b17136
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Apr 18 00:04:28 2012 +0200
libutil: let opendisk(3) try /dev
If a device node is given without path, and opening the node fails
initially, prepend "/dev/" to the node name and try opening again.
This is more in line with NetBSD behavior.
commit 4b6bf57bc77973874b041239c5a7809448e4d857
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Sep 10 23:57:32 2013 +0200
Block drivers: make IOCTL request unsigned long
The block driver protocol and libblockdriver's bdr_ioctl hook are
changed, as well as the users of this hook. Other parts of the system
are expected to change accordingly eventually, since the ioctl(2)
prototype has been aligned with NetBSD's.
Change-Id: Ide46245b22cfa89ed267a38088fb0ab7696eba92
commit d74ea98bdfd90605dd5b0c57e12390dc4dba56e6
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Jul 27 00:49:49 2013 +0200
Block protocol: add user endpoint to IOCTL request
I/O control requests now come with the endpoint of the user process
that initiated the ioctl(2) call. It is stored in a new BDEV_USER
field, which is an alias for BDEV_FLAGS. The contents of this field
are to be used only in highly specific situations. It should be
preserved (not replaced!) by services that forward IOCTL requests,
and may be set to NONE for service-initiated IOCTL requests.
Change-Id: I68a01b9ce43eca00e61b985a9cf87f55ba683de4
commit 15afe7ddf85235f00e5a5d0e825d2894ac58549d
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Jul 27 00:49:49 2013 +0200
Block protocol: use own [RW]_BIT definitions
The original R_BIT and W_BIT definitions have nothing to do with the
way these bits are used. Their distinct usage is more apparent when
they have different names.
Change-Id: Ia984457f900078b2e3502ceed565fead4e5bb965
commit c4c9435b0df31b36c11f27370475694043c5b8e0
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Sep 10 23:35:15 2013 +0200
libblockdriver: expose BLOCKDRIVER_MAX_DEVICES
This constant determines the range of valid device_id_t values that
a block driver can return from the bdr_device hook: a value between
0 and (BLOCKDRIVER_MAX_DEVICES - 1) inclusive.
Change-Id: I80fac469e88ac13d4b869007e6f2c2f7569da433
commit 2c5bd1565b5d9b487899dee7a023a641343e65a9
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Sep 11 13:33:00 2013 +0200
libblockdriver: various updates
- internal structure rearrangement;
- respond to char device open requests to avoid hanging VFS threads;
- make drivers use designated initializers;
- use devminor_t for all minor device numbers;
- change bdr_other hook to take ipc_status and return nothing;
- fix default geometry computation;
- add support for sef_cancel.
Change-Id: Ia063a136a3ddb2b78de36180feda870605753d70
commit b1cfc3123b03d5a8c9452ae251892593dc328357
Author: David van Moolenbroek <david@minix3.org>
Date: Sun Sep 1 14:34:17 2013 +0200
Block drivers: reply ENOTTY to unknown IOCTLs
Change-Id: Ie2e82d2491d546f4dd73b009100646e249a147b5
commit b5eab0554511ace3d7fae40883be0da139b0e2ab
Author: David van Moolenbroek <david@minix3.org>
Date: Sun Jul 28 14:03:07 2013 +0200
Move SUB_PER_DRIVE definition into minix/drvlib.h
commit 20c8aedc6737fc935ccea0b394616851ff4fdece
Author: David van Moolenbroek <david@minix3.org>
Date: Mon Sep 2 17:34:44 2013 +0200
VFS: update filp_pos on chardev I/O (workaround)
Previously, reading from or writing to a character device would not
update the file position on the corresponding filp object. Performing
this update correctly is not trivial: during and after the I/O
operation, the filp object must not be locked. Ideally, read/write
requests on a filp that is already involved in a read/write operation,
should be queued. For now, we optimistically update the file position
at the start of the I/O; this works under the assumptions listed in
the corresponding comment.
Change-Id: I172a61781850423709924390ae3df1f2d1f94707
commit dbef7bae2308a2e4449b8f274898ed3b60b2add2
Author: David van Moolenbroek <david@minix3.org>
Date: Mon Sep 2 13:45:02 2013 +0200
I2C: change BUSC_I2C_xxx to use own protocol
Previously it would use bits of the character driver protocol, which
will change heavily. In the new situation, the BUSC_I2C_xxx requests
use a protocol more in line with the PCI protocol, with the reply code
in m_type.
Change-Id: I51597b3f191078c8178ce17372de123031f7a4c4
commit 7cdb9683ec718247b03cbae3c827d9834c484e91
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Aug 31 16:13:37 2013 +0200
tests: add test77 for opening/closing PTYs
Change-Id: I30e3418f75137aa08037fa6581ff3d4cce32a114
commit ed9ee8b9cd5a647130a81bd70af48cee4f6dca15
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Aug 31 16:08:25 2013 +0200
TTY: fix for PTY open/close logic
Opening and closing the master side of a pseudo terminal without
opening the slave side would result in the pseudo terminal becoming
permanently unavailable. In addition, reopening the slave side
would be possible but not allow for I/O. Finally, attempting to
open an in-use master would wipe its I/O state. These issues have
been resolved.
Change-Id: I9235e3d9aba321803f9280b86b6b5e3646ad5ef3
commit cd14d52a99df7fb6dd7c6c48a034e4bf7c170789
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Aug 30 18:48:56 2013 +0200
tests: remove select subdirectory
This test set has been obsoleted by test40.
Change-Id: I55439152824906778ad07d409dfb327ac10bea70
commit 5ab6f6e7c2f4618648f163af6bd705e5ef2fafb8
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Aug 30 18:43:23 2013 +0200
tests: add test76 for interrupting VFS operations
Change-Id: Ic436cac61de8c42e0c7ee2d442c647528654cde9
commit ad548d2c9f78e2f2dc2108c50b6c7ed6755a53f1
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Aug 28 13:08:16 2013 +0200
VFS: fix interruption of blocking pipe operations
POSIX states that when interrupted, partially successful pipe
operations should return the partial result rather than EINTR. VFS
previously wouldn't look at the partial result, and not clear it
either, which would result in a panic upon the next pipe operation.
Change-Id: Ia1eb72b4b77394051444e63a1390d49bb315eb04
commit 813198ce7b0ce171f79529e8f6c5a94ab26ce1cb
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Aug 30 00:57:16 2013 +0200
libmthread: do not dump stack for free threads
Change-Id: Ic438a252f5bddaf1513f554c71173e6fffb0c674
commit 53b6001312428a1ac7379718616cb640614f53e6
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Aug 30 14:00:50 2013 +0200
VFS: worker thread model overhaul
The main purpose of this patch is to fix handling of unpause calls
from PM while another call is ongoing. The solution to this problem
sparked a full revision of the threading model, consisting of a large
number of related changes:
- all active worker threads are now always associated with a process,
and every process has at most one active thread working for it;
- the process lock is always held by a process's worker thread;
- a process can now have both normal work and postponed PM work
associated to it;
- timer expiry and non-postponed PM work is done from the main thread;
- filp garbage collection is done from a thread associated with VFS;
- reboot calls from PM are now done from a thread associated with PM;
- the DS events handler is protected from starting multiple threads;
- support for a system worker thread has been removed;
- the deadlock recovery thread has been replaced by a parameter to the
worker_start() function; the number of worker threads has
consequently been increased by one;
- saving and restoring of global but per-thread variables is now
centralized in worker_suspend() and worker_resume(); err_code is now
saved and restored in all cases;
- the concept of jobs has been removed, and job_m_in now points to a
message stored in the worker thread structure instead;
- the PM lock has been removed;
- the separate exec lock has been replaced by a lock on the VM
process, which was already being locked for exec calls anyway;
- PM_UNPAUSE is now processed as a postponed PM request, from a thread
associated with the target process;
- the FP_DROP_WORK flag has been removed, since it is no longer more
than just an optimization and only applied to processes operating on
a pipe when getting killed;
- assignment to "fp" now takes place only when obtaining new work in
the main thread or a worker thread, when resuming execution of a
thread, and in the special case of exiting processes during reboot;
- there are no longer special cases where the yield() call is used to
force a thread to run.
Change-Id: I7a97b9b95c2450454a9b5318dfa0e6150d4e6858
commit 6fbb7f4167510fb1a71fdad2b965dc2f7417aef4
Author: David van Moolenbroek <david@minix3.org>
Date: Wed Aug 28 15:58:30 2013 +0200
Retire ptrace(T_DUMPCORE), dumpcore(1), gcore(1)
The T_DUMPCORE implementation was not only broken - it would currently
produce a coredump of the tracer process rather than the traced
process - but also deeply flawed, and fixing it would require serious
alteration of PM's internal state machine. It should be possible to
implement the same functionality in userland, and that is now the
suggested way forward. For now, also remove the (identical) utilities
using T_DUMPCORE: dumpcore(1) and gcore(1).
Change-Id: I1d51be19c739362b8a5833de949b76382a1edbcc
commit fde44a0523a26b5bad5a0e3fc2921b1b5821d5d4
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Aug 30 13:42:51 2013 +0200
VFS: process char driver replies from main thread
Previously, processing of some replies coming from character drivers
could block on locks, and therefore, such processing was done from
threads that were associated to the character driver process. The
hidden consequence of this was that if all threads were in use, VFS
could drop replies coming from the driver. This patch returns VFS to
a situation where the replies from character drivers are processed
instantly from the main thread, by removing the situations that may
cause VFS to block while handling those replies.
- change the locking model for select, so that it will never block
on any processing that happens after the select call has been set
up, in particular processing of character driver select replies;
- clearly mark all select routines that may never block;
- protect against race conditions in do_select as result of the
locking that still does happen there (as is required for pipes);
- also handle select timers from the main thread;
- move processing of character driver replies into device.c.
Change-Id: I4dc8e69f265cbd178de0fbf321d35f58f067cc57
commit d5d67d293d4a7a2669eb179621a1c6d3e2796d97
Author: David van Moolenbroek <david@minix3.org>
Date: Sun Aug 25 00:26:38 2013 +0200
VFS: properly cancel select queries on unpause
Change-Id: I16e71db3f5c1bcc7ba6045bc9f02b13d71dc31eb
commit d0371a4fbdaddbf7bc543a8046bb91454e09283d
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Aug 30 13:33:56 2013 +0200
VFS: remove support for sync char driver protocol
Change-Id: I57cc870a053b813b3a3fc45da46606ea84fe4cb1
commit 2b12ad511fabd5a1019f736221f2940bfe7e2781
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Aug 30 13:00:44 2013 +0200
VFS: remove FP_BLOCKED_ON_DOPEN
These days, DEV_OPEN calls to character drivers block the calling
thread until completion or failure, and thus never return SUSPEND to
the caller. The same already applied to BDEV_OPEN calls to block
drivers. It has thus become impossible for a process to enter a state
of being blocked on a device open call.
There is currently no support for restarting device open calls to
restarted character drivers. This support was present in the _DOPEN
logic, but was already no longer triggering. In the future, this case
should be handled by the thread performing the open request.
Change-Id: I6cc1e7b4c9ed116c6ce160b315e6e060124dce00
commit 0eee377a337f5800760d2f5611d203c158ac6099
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Aug 30 12:18:27 2013 +0200
PFS: remember request information for IOCTLs
Not doing so caused PFS to commit protocol violations by relying on
stale information when sending replies. This stale information always
happened to be correct, which is why the problem went unnoticed.
Change-Id: Ia42ca670718d6e731193cd2c34a3ff455f8a94d3
commit 3272c645aca015b6778a8fe945e014a0a5acc023
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Aug 30 11:14:03 2013 +0200
Retire the synchronous character driver protocol
- change all sync char drivers into async drivers;
- retire support for the sync protocol in libchardev;
- remove async dev style, as this is now the default;
- remove dev_status from VFS;
- clean up now-unused protocol messages.
Change-Id: I6aacff712292f6b29f2ccd51bc1e7d7003723e87
commit 831d825d5a6891b79e97002e1a2d396dcb47bea6
Author: David van Moolenbroek <david@minix3.org>
Date: Fri Aug 30 10:48:34 2013 +0200
Sync char protocol: add nonblocking transfer flag
The async char protocol already has this, so this patch closes the
gap between the two protocols a bit. Support for this flag has been
added to all sync char drivers that support CANCEL at all.
The LOG driver was already using the asynchronous protocol, but it
did not support the nonblocking transfer flag. This has been fixed
as well.
Change-Id: Ia55432c9f102765b59ad3feb45a8bd47a782c93f
commit 4f5aec7328128b9d035ff1158f0e6a0f7bd393e4
Author: David van Moolenbroek <david@minix3.org>
Date: Sat Aug 24 12:29:39 2013 +0200
VFS: set w_drv_sendrec only when needed
As with w_task, this ensures that the field remains cleared if it is
not used. Without this, worker_stop could mistakenly identify a thread
as talking to a device driver rather than a (crashed) file server.
Change-Id: I7d3ebed3efc3cd4f5c891f61c67a6463109b6376
commit 9bbf923154733f1b9ba52a063872ccdbaa7ff0f0
Author: Thomas Veerman <thomas@minix3.org>
Date: Sat Aug 24 12:23:41 2013 +0200
VFS: set w_task only when needed
It was always set, but not always cleared, when talking to asynchronous
drivers. This could cause erratic behavior upon a driver crash.
Normally, a worker thread's w_task field is set when it's about to
communicate with a driver or FS. Then upon receiving a reply we can
do sanity checks (that the thread we want to wake up was actually
waiting for a reply). Also, when a driver/FS crashes, we can identify
which worker threads were talking to the crashed endpoint and handle
the error gracefully.
Asynchronous drivers are a bit special, though. In most cases, the
sender of the request is not interested in the reply (the sender was
suspended and only wants to know whether the request was successfully
caried out or not). However, the open request is special, as the reply
carries information needed by the sender. This is the only request
where a worker thread actually yields and waits for the result. This is
also the only case where we're interested in setting w_task for
asynchronous drivers.
Change-Id: Ia1ce2747937df376122b5e13b6a069de27fcc379
commit 2a293ca6703354cd7be8eb18147f7c4d18894132
Author: David van Moolenbroek <david@minix3.org>
Date: Mon Aug 19 20:34:15 2013 +0200
Import NetBSD df(1)
Change-Id: Ia60f8b23b961e4132efece1e7e38f0d63597c13b
commit c35fc992ccc052e1f3741be0db9b676682f68bf8
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Aug 20 01:33:43 2013 +0200
Enable getmntinfo(3)
Change-Id: Id9d11a67cdcad331a52030127f7c7ead3c847919
commit 907b40860983a73fe22f45aa8b175b27cb7dc1d7
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Aug 20 00:49:07 2013 +0200
test55: add tests for getvfsstat(2)
Change-Id: Iad4567068a82bb0e780891a478ba5a06b63f1d48
commit 3297eb8d103545690e9f2cb6c63fdd1ce3452c6d
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Aug 20 01:28:23 2013 +0200
Implement support for [f]statvfs1(2)
The [f]statvfs(3) calls now use [f]statvfs1(2).
Change-Id: I56c92fe1c70670f476631c4a50e47351562712c6
commit 4f576aff9f046276974c4a769427a1cb8c2470a0
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Aug 20 01:39:47 2013 +0200
Implement support for getvfsstat(2)
Change-Id: I4680c071b94fa855edfb6220b03cff6b20f3c1c9
commit 53485a9b35ec2e27ec0df232d8c80db1a829246b
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Aug 20 01:37:18 2013 +0200
Redo mount(2)/umount(2) ABI
- pass in file system type through mount(2), and return this type in
statvfs structures as generated by [f]statvfs(2);
- align mount flags field with NetBSD's, splitting out service flags
which are not to be passed to VFS;
- remove limitation of mount ABI to 16-byte labels, so that labels
can be made larger in the future;
- introduce new m11 message union type for mount(2) as side effect.
Change-Id: Ia9ed4566d88af5239749c57274312be26bb61448
commit 6ebff29307cd1bacba965d1e8dfb465b6684ccb7
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Aug 20 01:35:35 2013 +0200
Align "struct statvfs" with NetBSD
This is a requirement for implementing calls such as getmntinfo(3).
VFS is now responsible for filling in some of the structure's fields.
Change-Id: I457b463c70ea2f326b9c50de954a91cfefba0065
commit 9b5871fd21d37b0533b6546c408ff4ad4ae02952
Author: David van Moolenbroek <david@minix3.org>
Date: Tue Aug 20 00:55:49 2013 +0200
VFS/FS: remove fstatfs(2) and REQ_FSTATFS
The fstatfs(3) call now uses fstatvfs(2).
Change-Id: Ic6bbb040e1a8f01aaffe1a6a809dd4c7a2eae2f6
Change-Id: I44109a83fd9155f15a8b580e9471ed1aff952673
700 lines
45 KiB
Plaintext
700 lines
45 KiB
Plaintext
## Description of VFS Thomas Veerman 21-3-2013
|
|
## This file is organized such that it can be read both in a Wiki and on
|
|
## the MINIX terminal using e.g. vi or less. Please, keep the file in the
|
|
## source tree as the canonical version and copy changes into the Wiki.
|
|
#pragma section-numbers 2
|
|
|
|
= VFS internals =
|
|
|
|
<<TableOfContents(2)>>
|
|
|
|
## Table of contents
|
|
## 1 ..... General description of responsibilities
|
|
## 2 ..... General architecture
|
|
## 3 ..... Worker threads
|
|
## 4 ..... Locking
|
|
## 4.1 .... Locking requirements
|
|
## 4.2 .... Three-level Lock
|
|
## 4.3 .... Data structures subject to locking
|
|
## 4.4 .... Locking order
|
|
## 4.5 .... Vmnt (file system) locking
|
|
## 4.6 .... Vnode (open file) locking
|
|
## 4.7 .... Filp (file position) locking
|
|
## 4.8 .... Lock characteristics per request type
|
|
## 5 ..... Recovery from driver crashes
|
|
## 5.1 .... Recovery from block drivers crashes
|
|
## 5.2 .... Recovery from character driver crashes
|
|
## 5.3 .... Recovery from File Server crashes
|
|
|
|
== General description of responsibilities ==
|
|
## 1 General description of responsibilities
|
|
VFS implements the file system in cooperation with one or more File Servers
|
|
(FS). The File Servers take care of the actual file system on a partition. That
|
|
is, they interpret the data structure on disk, write and read data to/from
|
|
disk, etc. VFS sits on top of those File Servers and communicates with
|
|
them. Looking inside VFS, we can identify several roles. First, a role of VFS
|
|
is to handle most POSIX system calls that are supported by Minix. Additionally,
|
|
it supports a few calls necessary for libc. The following system calls are
|
|
handled by VFS:
|
|
|
|
access, chdir, chmod, chown, chroot, close, creat, fchdir, fcntl, fstat,
|
|
fstatvfs, fsync, ftruncate, getdents, getvfsstat, ioctl, link, llseek, lseek,
|
|
lstat, mkdir, mknod, mount, open, pipe, read, readlink, rename, rmdir, select,
|
|
stat, statvfs, symlink, sync, truncate, umask, umount, unlink, utime, write.
|
|
|
|
Second, it maintains part of the state belonging to a process (process state is
|
|
spread out over the kernel, VM, PM, and VFS). For example, it maintains state
|
|
for select(2) calls, file descriptors and file positions. Also, it cooperates
|
|
with the Process Manager to handle the fork, exec, and exit system calls.
|
|
Third, VFS keeps track of endpoints that are supposed to be drivers for
|
|
character or block special files. File Servers can be regarded as drivers for
|
|
block special files, although they are handled entirely different compared
|
|
to other drivers.
|
|
|
|
The following diagram depicts how a read() on a file in /home is being handled:
|
|
{{{
|
|
----------------
|
|
| user process |
|
|
----------------
|
|
^ ^
|
|
| |
|
|
read(2) \
|
|
| \
|
|
V \
|
|
---------------- |
|
|
| VFS | |
|
|
---------------- |
|
|
^ |
|
|
| |
|
|
V |
|
|
------- -------- ---------
|
|
| MFS | | MFS | | MFS |
|
|
| / | | /usr | | /home |
|
|
------- -------- ---------
|
|
}}}
|
|
Diagram 1: handling of read(2) system call
|
|
|
|
The user process executes the read system call which is delivered to VFS. VFS
|
|
verifies the read is done on a valid (open) file and forwards the request
|
|
to the FS responsible for the file system on which the file resides. The FS
|
|
reads the data, copies it directly to the user process, and replies to VFS
|
|
it has executed the request. Subsequently, VFS replies to the user process
|
|
the operation is done and the user process continues to run.
|
|
|
|
== General architecture ==
|
|
## 2 General architecture
|
|
VFS works roughly identical to every other server and driver in Minix; it
|
|
fetches a message (internally referred to as a job in some cases), executes
|
|
the request embedded in the message, returns a reply, and fetches the next
|
|
job. There are several sources for new jobs: from user processes, from PM, from
|
|
the kernel, and from suspended jobs inside VFS itself (suspended operations
|
|
on pipes, locks, or character special files). File Servers are regarded as
|
|
normal user processes in this case, but their abilities are limited. This
|
|
is to prevent deadlocks. Once a job is received, a worker thread starts
|
|
executing it. During the lifetime of a job, the worker thread might need
|
|
to talk to several File Servers. The protocol VFS speaks with File Servers
|
|
is fully documented on the Wiki at [0]. The protocol fields are defined in
|
|
<minix/vfsif.h>. If the job is an operation on a character or block special
|
|
file and the need to talk to a driver arises, VFS uses the Character and
|
|
Block Device Protocol. See [1]. This is sadly not official documentation,
|
|
but it is an accurate description of how it works. Luckily, driver writers
|
|
can use the libchardriver and libblockdriver libraries and don't have to
|
|
know the details of the protocol.
|
|
|
|
== Worker threads ==
|
|
## 3 Worker threads
|
|
Upon start up, VFS spawns a configurable amount of worker threads. The
|
|
main thread fetches requests and replies, and hands them off to idle or
|
|
reply-pending workers, respectively. If no worker threads are available,
|
|
the request is queued. All standard system calls are handled by such worker
|
|
threads. One of the threads is reserved to handle new requests from system
|
|
processes (i.e., File Servers and drivers) when there are no normal worker
|
|
threads available; all normal threads might be blocked on a single worker
|
|
thread that caused a system process to send a request on its own. To unblock
|
|
all normal threads, we need to reserve one spare thread to handle that
|
|
situation. VFS drives all File Servers and drivers asynchronously. While
|
|
waiting for a reply, a worker thread is blocked and other workers can keep
|
|
processing requests. Upon reply the worker thread is unblocked.
|
|
|
|
As mentioned above, the main thread is responsible for retrieving new jobs and
|
|
replies to current jobs and start or unblock the proper worker thread.
|
|
Driver replies are processed directly from the main thread. As a consequence,
|
|
these processing routines may not block their calling thread. In some cases,
|
|
these routines may resume a thread that is blocked waiting for the reply. This
|
|
is always the case for block driver replies, and may or may not be the case for
|
|
character driver replies. The character driver reply processing routines may
|
|
also unblock suspended processes which in turn generate new jobs to be handled
|
|
by the main loop (e.g., suspended reads and writes on pipes). So depending
|
|
on the reply a new thread may have to be started.
|
|
|
|
Worker threads are strictly tied to a process, and each process can have at
|
|
most one worker thread running for it. Generally speaking, there are two types
|
|
of work supported by worker threads: normal work, and work from PM. The main
|
|
subtype of normal work is the handling of a system call made by the process
|
|
itself. The process is blocked while VFS is handling the system call, so no new
|
|
system call can arrive from a process while VFS has not completed a previous
|
|
system call from that process. For that reason, if there are no worker threads
|
|
available to handle the work, the work is queued in the corresponding process
|
|
entry of the fproc table.
|
|
|
|
The other main type of work consists of requests from PM. The protocol PM
|
|
speaks with VFS is asynchronous. PM is allowed to send up to one request per
|
|
process to VFS, in addition to a request to initiate a reboot. Most jobs from
|
|
PM are taken care of immediately by the main thread, but some jobs require a
|
|
worker thread context (to be able to sleep) and/or serialization with normal
|
|
work. Therefore, each process may have a PM request queued for execution, also
|
|
in the fproc table. Managing proper queuing, addition, and execution of both
|
|
normal and PM work is the responsibility of the worker thread infrastructure.
|
|
|
|
There are several special tasks that require a worker thread, and these are
|
|
implemented as normal work associated with a certain special process that does
|
|
not make regular VFS calls anyway. For example, the initial ramdisk mount
|
|
procedure uses a thread associated with the VFS process. Some of these special
|
|
tasks require protection against being started multiple times at once, as this
|
|
is not only undesirable but also disallowed. The full list of worker thread
|
|
task types and subtypes is shown in Table 1.
|
|
|
|
{{{
|
|
-------------------------------------------------------------------------
|
|
| Worker thread task | Type | Association | May use spare? |
|
|
+---------------------------+--------+-----------------+----------------+
|
|
| system call from process | normal | calling process | if system proc |
|
|
+---------------------------+--------+-----------------+----------------+
|
|
| resumed pipe operation | normal | calling process | no |
|
|
+---------------------------+--------+-----------------+----------------+
|
|
| postponed PM request | PM | target process | no |
|
|
+---------------------------+--------+-----------------+----------------+
|
|
| DS event notification | normal | DS | yes |
|
|
+---------------------------+--------+-----------------+----------------+
|
|
| initial ramdisk mounting | normal | VFS | no |
|
|
+---------------------------+--------+-----------------+----------------+
|
|
| reboot sequence | normal | PM | no |
|
|
-------------------------------------------------------------------------
|
|
}}}
|
|
Table 1: worker thread work types and subtypes
|
|
|
|
Communication with block drivers is asynchronous, but at this time, access to
|
|
these drivers is serialized on a per-driver basis. File Servers are treated
|
|
differently. VFS was designed to be able to send requests concurrently to File
|
|
Servers, although at the time of writing there are no File Servers that can
|
|
actually make use of that functionality. To identify which reply from an FS
|
|
belongs to which worker thread, all requests have an embedded transaction
|
|
identification number (a magic number + thread id encoded in the mtype field of
|
|
a message) which the FS has to echo upon reply. Because the range of valid
|
|
transaction IDs is isolated from valid system call numbers, VFS can use that ID
|
|
to differentiate between replies from File Servers and actual new system calls
|
|
from FSes. Using this mechanism VFS is able to support FUSE and ProcFS.
|
|
|
|
== Locking ==
|
|
## 4 Locking
|
|
To ensure correct execution of system calls, worker threads sometimes need
|
|
certain objects within VFS to remain unchanged during thread suspension
|
|
and resumption (i.e., when they need to communicate with a driver or File
|
|
Server). Threads keep most state on the stack, but there are a few global
|
|
variables that require protection: the fproc table, vmnt table, vnode table,
|
|
and filp table. Other tables such as lock table, select table, and dmap table
|
|
don't require protection by means of exclusive access. There it's required
|
|
and enough to simply mark an entry in use.
|
|
|
|
=== Locking requirements ===
|
|
## 4.1 Locking requirements
|
|
VFS implements the locking model described in [2]. For completeness of this
|
|
document we'll describe it here, too. The requirements are based on a threading
|
|
package that is non-preemptive. VFS must guarantee correct functioning with
|
|
several, semi-concurrently executing threads in any arbitrary order. The
|
|
latter requirement follows from the fact that threads need service from
|
|
other components like File Servers and drivers, and they may take any time
|
|
to complete requests.
|
|
1. Consistency of replicated values. Several system calls rely on VFS keeping a replicated representation of data in File Servers (e.g., file sizes, file modes, etc.).
|
|
1. Isolation of system calls. Many system calls involve multiple requests to FSes. Concurrent requests from other processes must not lead to otherwise impossible results (e.g., a chmod operation on a file cannot fail halfway through because it's suddenly unlinked or moved).
|
|
1. Integrity of objects. From the point of view of threads, obtaining mutual exclusion is a potentially blocking operation. The integrity of any objects used across blocking calls must be guaranteed (e.g., the file mode in a vnode must remain intact not only when talking to other components, but also when obtaining a lock on a filp).
|
|
1. No deadlock. Not one call may cause another call to never complete. Deadlock situations are typically the result of two or more threads that each hold exclusive access to one resource and want exclusive access to the resource held by the other thread. These resources are a) data (global variables) and b) worker threads.
|
|
a. Conflicts between locking of different types of objects can be avoided by keeping a locking order: objects of different type must always be locked in the same order. If multiple objects of the same type are to be locked, then first a "common denominator" higher up in the locking order must be locked.
|
|
a. Some threads can only run to completion when another thread does work on their behalf. Examples of this are drivers and file servers that do system calls on their own (e.g., ProcFS, PFS/UNIX Domain Sockets, FUSE) or crashing components (e.g., a driver for a character special file that crashes during a request; a second thread is required to handle resource clean up or driver restart before the first thread can abort or retry the request).
|
|
1. No starvation. VFS must guarantee that every system call completes in finite time (e.g., an infinite stream of reads must never completely block writes). Furthermore, we want to maximize parallelism to improve performance. This leads to:
|
|
1. A request to one File Server must not block access to other FS processes. This means that most forms of locking cannot take place at a global level, and must at most take place on the file system level.
|
|
1. No read-only operation on a regular file must block an independent read call to that file. In particular, (read-only) open and close operations may not block such reads, and multiple independent reads on the same file must be able to take place concurrently (i.e., reads that do not share a file position between their file descriptors).
|
|
|
|
=== Three-level Lock ===
|
|
## 4.2 Three-level Lock
|
|
From the requirements it follows that we need at least two locking types: read
|
|
and write locks. Concurrent reads are allowed, but writes are exclusive both
|
|
from reads and from each other. However, in a lot of cases it possible to use
|
|
a third locking type that is in between read and write lock: the serialize
|
|
lock. This is implemented in the three-level lock [2]. The three-level
|
|
lock provides:
|
|
TLL_READ: allows an unlimited number of threads to hold the lock with the
|
|
same type (both the thread itself and other threads); N * concurrent.
|
|
TLL_READSER: also allows an unlimited number of threads with type TLL_READ,
|
|
but only one thread can obtain serial access to the lock; N * concurrent +
|
|
1 * serial.
|
|
TLL_WRITE: provides full mutual exclusion; 1 * exclusive + 0 * concurrent +
|
|
0 * serial.
|
|
In absence of TLL_READ locks, a TLL_READSER is identical to TLL_WRITE. However,
|
|
TLL_READSER never blocks concurrent TLL_READ access. TLL_READSER can be
|
|
upgraded to TLL_WRITE; the thread will block until the last TLL_READ lock
|
|
leaves and new TLL_READ locks are blocked. Locks can be downgraded to a
|
|
lower type. The three-level lock is implemented using two FIFO queues with
|
|
write-bias. This guarantees no starvation.
|
|
|
|
=== Data structures subject to locking ===
|
|
## 4.3 Data structures subject to locking
|
|
VFS has a number of global data structures. See Table 2.
|
|
{{{
|
|
--------------------------------------------------------------------
|
|
| Structure | Object description |
|
|
+------------+-----------------------------------------------------|
|
|
| fproc | Process (includes process's file descriptors) |
|
|
+------------+-----------------------------------------------------|
|
|
| vmnt | Virtual mount; a mounted file system |
|
|
+------------+-----------------------------------------------------|
|
|
| vnode | Virtual node; an open file |
|
|
+------------+-----------------------------------------------------|
|
|
| filp | File position into an open file |
|
|
+------------+-----------------------------------------------------|
|
|
| lock | File region locking state for an open file |
|
|
+------------+-----------------------------------------------------|
|
|
| select | State for an in-progress select(2) call |
|
|
+------------+-----------------------------------------------------|
|
|
| dmap | Mapping from major device number to a device driver |
|
|
--------------------------------------------------------------------
|
|
}}}
|
|
Table 2: VFS object types.
|
|
|
|
An fproc object is a process. An fproc object is created by fork(2)
|
|
and destroyed by exit(2) (which may, or may not, be instantiated from the
|
|
process itself). It is identified by its endpoint number ('fp_endpoint')
|
|
and process id ('fp_pid'). Both are unique although in general the endpoint
|
|
number is used throughout the system.
|
|
A vmnt object is a mounted file system. It is created by mount(2) and destroyed
|
|
by umount(2). It is identified by a device number ('m_dev') and FS endpoint
|
|
number ('m_fs_e'); both are unique to each vmnt object. There is always a
|
|
single process that handles a file system on a device and a device cannot
|
|
be mounted twice.
|
|
A vnode object is the VFS representation of an open inode on the file
|
|
system. A vnode object is created when a first process opens or creates the
|
|
corresponding file and is destroyed when the last process, which has that
|
|
file open, closes it. It is identified by a combination of FS endpoint number
|
|
('v_fs_e') and inode number of that file system ('v_inode_nr'). A vnode
|
|
might be mapped to another file system; the actual reading and writing is
|
|
handled by a different endpoint. This has no effect on locking.
|
|
A filp object contains a file position within a file. It is created when a file
|
|
is opened or anonymous pipe created and destroyed when the last user (i.e.,
|
|
process) closes it. A file descriptor always points to a single filp. A filp
|
|
always point to a single vnode, although not all vnodes are pointed to by a
|
|
filp. A filp has a reference count ('filp_count') which is identical to the
|
|
number of file descriptors pointing to it. It can be increased by a dup(2)
|
|
or fork(2). A filp can therefore be shared by multiple processes.
|
|
A lock object keeps information about locking of file regions. This has
|
|
nothing to do with the threading type of locking. The lock objects require
|
|
no locking protection and won't be discussed further.
|
|
A select object keeps information on a select(2) operation that cannot
|
|
be fulfilled immediately (waiting for timeout or file descriptors not
|
|
ready). They are identified by their owner ('requestor'); a pointer to the
|
|
fproc table. A null pointer means not in use. A select object can be used by
|
|
only one process and a process can do only one select(2) at a time. Select(2)
|
|
operates on filps and is organized in such a way that it is sufficient to
|
|
apply locking on individual filps and not on select objects themselves. They
|
|
won't be discussed further.
|
|
A dmap object is a mapping from a device number to a device driver. A device
|
|
driver can have multiple device numbers associated (e.g., TTY). Access to
|
|
a driver is exclusive when it uses the synchronous driver protocol.
|
|
|
|
=== Locking order ===
|
|
## 4.4 Locking order
|
|
Based on the description in the previous section, we need protection for
|
|
fproc, vmnt, vnode, and filp objects. To prevent deadlocks as a result of
|
|
object locking, we need to define a strict locking order. In VFS we use the
|
|
following order:
|
|
|
|
{{{
|
|
fproc > [exec] > vmnt > vnode > filp > [block special file] > [dmap]
|
|
}}}
|
|
|
|
That is, no thread may lock an fproc object while holding a vmnt lock,
|
|
and no thread may lock a vmnt object while holding an (associated) vnode, etc.
|
|
|
|
Fproc needs protection because processes themselves can initiate system
|
|
calls, but also PM can cause system calls that have to be executed in their
|
|
name. For example, a process might be busy reading from a character device
|
|
and another process sends a termination signal. The exit(2) that follows is
|
|
sent by PM and is to be executed by the to-be-killed process itself. At this
|
|
point there is contention for the fproc object that belongs to the process,
|
|
hence the need for protection. This problem is solved in a simple way. Recall
|
|
that all worker threads are bound to a process. This also forms the basis of
|
|
fproc locking: each worker thread acquires and holds the fproc lock for its
|
|
associated process for as long as it is processing work for that process.
|
|
|
|
There are two cases where a worker thread may hold the lock to more than one
|
|
process. First, as mentioned, the reboot procedure is executed from a worker
|
|
thread set in the context of the PM process, thus with the PM process entry
|
|
lock held. The procedure itself then acquires a temporary lock on every other
|
|
process in turn, in order to clean it up without interference. Thus, the PM
|
|
process entry is higher up in the locking order than all other process entries.
|
|
|
|
Second, the exec(2) call is protected by a lock, and this exec lock is
|
|
currently implemented as a lock on the VM process entry. The exec lock is
|
|
acquired by a worker thread for the process performing the exec(2) call, and
|
|
thus, the VM process entry is below all other process entries in the locking
|
|
order. The exec(2) call is protected by a lock for the following reason. VFS
|
|
uses a number of variables on the heap to read ELF headers. They are on the
|
|
heap due to their size; putting them on the stack would increase stack size
|
|
demands for worker threads. The exec call does blocking read calls and thus
|
|
needs exclusive access to these variables. However, only the exec(2) syscall
|
|
needs this lock.
|
|
|
|
Access to block special files needs to be exclusive. File Servers are
|
|
responsible for handling reads from and writes to block special files; if
|
|
a block special file is on a device that is mounted, the FS responsible for
|
|
that mount point takes care of it, otherwise the FS that handles the root of
|
|
the file system is responsible. Due to mounting and unmounting file systems,
|
|
the FS handling a block special file may change. Locking the vnode is not
|
|
enough since the inode can be on an entirely different File Server. Therefore,
|
|
access to block special files must be mutually exclusive from concurrent
|
|
mount(2)/umount(2) operations. However, when we're not accessing a block
|
|
special file, we don't need this lock.
|
|
|
|
=== Vmnt (file system) locking ===
|
|
## 4.5 Vmnt (file system) locking
|
|
Vmnt locking cannot be seen completely separately from vnode locking. For
|
|
example, umount(2) fails if there are still in-use vnodes, which means that
|
|
FS requests [0] only involving in-use inodes do not have to acquire a vmnt
|
|
lock. On the other hand, all other request do need a vmnt lock. Extrapolating
|
|
this to system calls this means that all system calls involving a file
|
|
descriptor don't need a vmnt lock and all other system calls (that make FS
|
|
requests) do need a vmnt lock.
|
|
{{{
|
|
-------------------------------------------------------------------------------
|
|
| Category | System calls |
|
|
+-------------------+---------------------------------------------------------+
|
|
| System calls with | access, chdir, chmod, chown, chroot, creat, dumpcore+, |
|
|
| a path name | exec, link, lstat, mkdir, mknod, mount, open, readlink, |
|
|
| argument | rename, rmdir, stat, statvfs, symlink, truncate, umount,|
|
|
| | unlink, utime |
|
|
+-------------------+---------------------------------------------------------+
|
|
| System calls with | close, fchdir, fcntl, fstat, fstatvfs, ftruncate, |
|
|
| a file descriptor | getdents, ioctl, llseek, pipe, read, select, write |
|
|
| argument | |
|
|
+-------------------+---------------------------------------------------------+
|
|
| System calls with | fsync++, getvfsstat, sync, umask |
|
|
| other or no | |
|
|
| arguments | |
|
|
-------------------------------------------------------------------------------
|
|
}}}
|
|
Table 3: System call categories.
|
|
+ path name argument is implicit, the path name is "core.<pid>"
|
|
++ although fsync actually provides a file descriptor argument, it's only
|
|
used to find the vmnt and not to do any actual operations on
|
|
|
|
Before we describe what kind of vmnt locks VFS applies to system calls with a
|
|
path name or other arguments, we need to make some notes on path lookup. Path
|
|
lookups take arbitrary paths as input (relative and absolute). They can start
|
|
at any vmnt (based on root directory and working directory of the process doing
|
|
the lookup) and visit any file system in arbitrary order, possibly visiting
|
|
the same file system more than once. As such, VFS can never tell in advance
|
|
at which File Server a lookup will end. This has the following consequences:
|
|
* In the lookup procedure, only one vmnt must be locked at a time. When
|
|
moving from one vmnt to another, the first vmnt has to be unlocked before
|
|
acquiring the next lock to prevent deadlocks.
|
|
* The lookup procedure must lock each visited file system with TLL_READSER
|
|
and downgrade or upgrade to the lock type desired by the caller for the
|
|
destination file system (as VFS cannot know which file system is final). This
|
|
is to prevent deadlocks when a thread acquires a TLL_READSER on a vmnt and
|
|
another thread TLL_READ on the same vmnt. If the second thread is blocked
|
|
on the first thread due to it acquiring a lock on a vnode, the first thread
|
|
will be unable to upgrade a TLL_READSER lock to TLL_WRITE.
|
|
|
|
We use the following mapping for vmnt locks onto three-level lock types:
|
|
{{{
|
|
-------------------------------------------------------------------------------
|
|
| Lock type | Mapped to | Used for |
|
|
+------------+-------------+--------------------------------------------------+
|
|
| VMNT_READ | TLL_READ | Read-only operations and fully independent write |
|
|
| | | operations |
|
|
+------------+-------------+--------------------------------------------------+
|
|
| VMNT_WRITE | TLL_READSER | Independent create and modify operations |
|
|
+------------+-------------+--------------------------------------------------+
|
|
| VMNT_EXCL | TLL_WRITE | Delete and dependent write operations |
|
|
-------------------------------------------------------------------------------
|
|
}}}
|
|
Table 4: vmnt to tll lock mapping
|
|
|
|
The following table shows a sub-categorization of system calls without a
|
|
file descriptor argument, together with their locking types and motivation
|
|
as used by VFS.
|
|
{{{
|
|
-------------------------------------------------------------------------------
|
|
| Group | System calls | Lock type | Motivation |
|
|
+-------------+--------------+------------+-----------------------------------+
|
|
| File open | chdir, | VMNT_READ | These operations do not interfere |
|
|
| ops. | chroot, exec,| | with each other, as vnodes can be |
|
|
| (non-create)| open | | opened concurrently, and open |
|
|
| | | | operations do not affect |
|
|
| | | | replicated state. |
|
|
+-------------+--------------+------------+-----------------------------------+
|
|
| File create-| creat, | VMNT_EXCL | File create ops. require mutual |
|
|
| and-open | open(O_CREAT)| for create | exclusion from concurrent file |
|
|
| ops | | VMNT_WRITE | open ops. If the file already |
|
|
| | | for open | existed, the VMNT_WRITE lock that |
|
|
| | | | is necessary for the lookup is |
|
|
| | | | not upgraded |
|
|
+-------------+--------------+------------+-----------------------------------+
|
|
| File create-| pipe | VMNT_READ | These create nameless inodes |
|
|
| unique-and- | | | which cannot be opened by means |
|
|
| open ops. | | | of a path. Their creation |
|
|
| | | | therefore does not interfere with |
|
|
| | | | anything else |
|
|
+-------------+--------------+------------+-----------------------------------+
|
|
| File create-| mkdir, mknod,| VMNT_WRITE | These operations do not affect |
|
|
| only ops. | slink | | any VFS state, and can therefore |
|
|
| | | | take place concurrently with open |
|
|
| | | | operations |
|
|
+-------------+--------------+------------+-----------------------------------+
|
|
| File info | access, lstat| VMNT_READ | These operations do not interfere |
|
|
| retrieval or| readlink,stat| | with each other and do not modify |
|
|
| modification| utime | | replicated state |
|
|
+-------------+--------------+------------+-----------------------------------+
|
|
| File | chmod, chown,| VMNT_READ | These operations do not interfere |
|
|
| modification| truncate | | with each other. They do need |
|
|
| | | | exclusive access on the vnode |
|
|
| | | | level |
|
|
+-------------+--------------+------------+-----------------------------------+
|
|
| File link | link | VMNT_WRITE | Identical to file create-only |
|
|
| ops. | | | operations |
|
|
+-------------+--------------+------------+-----------------------------------+
|
|
| File unlink | rmdir, unlink| VMNT_EXCL | These must not interfere with |
|
|
| ops. | | | file create operations, to avoid |
|
|
| | | | the scenario where inodes are |
|
|
| | | | reused immediately. However, due |
|
|
| | | | to necessary path checks, the |
|
|
| | | | vmnt is first locked VMNT_WRITE |
|
|
| | | | and then upgraded |
|
|
+-------------+--------------+------------+-----------------------------------+
|
|
| File rename | rename | VMNT_EXCL | Identical to file unlink |
|
|
| ops. | | | operations |
|
|
+-------------+--------------+------------+-----------------------------------+
|
|
| Non-file | sync, umask, | VMNT_READ | umask does not involve the file |
|
|
| ops. | getvfsstat | or none | system, so it does not need |
|
|
| | | | locks. sync does not alter state |
|
|
| | | | in VFS and is atomic at the FS |
|
|
| | | | level. getvfsstat caches stats |
|
|
| | | | only and requires no exclusion. |
|
|
-------------------------------------------------------------------------------
|
|
}}}
|
|
Table 5: System call without file descriptor argument sub-categorization
|
|
|
|
=== Vnode (open file) locking ===
|
|
## 4.6 Vnode (open file) locking
|
|
Compared to vmnt locking, vnode locking is relatively straightforward. All
|
|
read-only accesses to vnodes that merely read the vnode object's fields are
|
|
allowed to be concurrent. Consequently, all accesses that change fields
|
|
of a vnode object must be exclusive. This leaves us with creation and
|
|
destruction of vnode objects (and related to that, their reference counts);
|
|
it's sufficient to serialize these accesses. This follows from the fact
|
|
that a vnode is only created when the first user opens it, and destroyed
|
|
when the last user closes it. A open file in process A cannot be be closed
|
|
by process B. Note that this also relies on the fact that a process can do
|
|
only one system call at a time. Kernel threads would violate this assumption.
|
|
|
|
We use the following mapping for vnode locks onto three-level lock types:
|
|
{{{
|
|
-------------------------------------------------------------------------------
|
|
| Lock type | Mapped to | Used for |
|
|
+------------+-------------+--------------------------------------------------+
|
|
| VNODE_READ | TLL_READ | Read access to previously opened vnodes |
|
|
+------------+-------------+--------------------------------------------------+
|
|
| VNODE_OPCL | TLL_READSER | Creation, opening, closing, and destruction of |
|
|
| | | vnodes |
|
|
+------------+-------------+--------------------------------------------------+
|
|
| VNODE_WRITE| TLL_WRITE | Write access to previously opened vnodes |
|
|
-------------------------------------------------------------------------------
|
|
}}}
|
|
Table 6: vnode to tll lock mapping
|
|
|
|
When vnodes are destroyed, they are initially locked with VNODE_OPCL. After
|
|
all, we're going to alter the reference count, so this must be serialized. If
|
|
the reference count then reaches zero we obtain exclusive access. This should
|
|
always be immediately possible unless there is a consistency problem. See
|
|
section 4.8 for an exhaustive listing of locking methods for all operations on
|
|
vnodes.
|
|
|
|
=== Filp (file position) locking ===
|
|
## 4.7 Filp (file position) locking
|
|
The main fields of a filp object that are shared between various processes
|
|
(and by extension threads), and that can change after object creation,
|
|
are filp_count and filp_pos. Writes to and reads from filp object must be
|
|
mutually exclusive, as all system calls have to use the latest version. For
|
|
example, a read(2) call changes the file position (i.e., filp_pos), so two
|
|
concurrent reads must obtain exclusive access. Consequently, as even read
|
|
operations require exclusive access, filp object don't use three-level locks,
|
|
but only mutexes.
|
|
|
|
System calls that involve a file descriptor often access both the filp and
|
|
the corresponding vnode. The locking order requires us to first lock the
|
|
vnode and then the filp. This is taken care of at the filp level. Whenever
|
|
a filp is locked, a lock on the vnode is acquired first. Conversely, when
|
|
a filp is unlocked, the corresponding vnode is also unlocked. A convenient
|
|
consequence is that whenever a vnode is locked exclusively (VNODE_WRITE),
|
|
all corresponding filps are implicitly locked. This is of particular use
|
|
when multiple filps must be locked at the same time:
|
|
* When opening a named pipe, VFS must make sure that there is at most one filp for the reader end and one filp for the writer end.
|
|
* Pipe readers and writers must be suspended in the absence of (respectively) writers and readers.
|
|
Because both filps are linked to the same vnode object (they are for the same
|
|
pipe), it suffices to exclusively lock that vnode instead of both filp objects.
|
|
|
|
In some cases it can happen that a function that operates on a locked filp,
|
|
calls another function that triggers another lock on a different filp for
|
|
the same vnode. For example, close_filp. At some point, close_filp() calls
|
|
release() which in turn will loop through the filp table looking for pipes
|
|
being select(2)ed on. If there are, the select code will lock the filp and do
|
|
operations on it. This works fine when doing a select(2) call, but conflicts
|
|
with close(2) or exit(2). Lock_filp() makes an exception for this situation;
|
|
if you've already locked a vnode with VNODE_OPCL or VNODE_WRITE when locking
|
|
a filp, you obtain a "soft lock" on the vnode for this filp. This means
|
|
that lock_filp won't actually try to lock the vnode (which wouldn't work),
|
|
but flags the vnode as "skip unlock_vnode upon unlock_filp." Upon unlocking
|
|
the filp, the vnode remains locked, the soft lock is removed, and the filp
|
|
mutex is released. Note that this scheme does not violate the locking order;
|
|
the vnode is (already) locked before the filp.
|
|
|
|
A similar problem arises with create_pipe. In this case we obtain a new vnode
|
|
object, lock it, and obtain two new, locked, filp objects. If everything works
|
|
out and the filp objects are linked to the same vnode, we run into trouble
|
|
when unlocking both filps. The first filp being unlocked would work; the
|
|
second filp doesn't have an associated vnode that's locked anymore. Therefore
|
|
we introduced a plural unlock_filps(filp1, filp2) that can unlock two filps
|
|
that both point to the same vnode.
|
|
|
|
=== Lock characteristics per request type ===
|
|
## 4.8 Lock characteristics per request type
|
|
For File Servers that support concurrent requests, it's useful to know which
|
|
locking guarantees VFS provides for vmnts and vnodes, so it can take that
|
|
into account when protecting internal data structures. READ = TLL_READ,
|
|
READSER = TLL_READSER, WRITE = TLL_WRITE. The vnode locks applies to the
|
|
REQ_INODE_NR field in requests, unless the notes say otherwise.
|
|
{{{
|
|
------------------------------------------------------------------------------
|
|
| request | vmnt | vnode | notes |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_BREAD | | READ | VFS serializes reads from and writes to |
|
|
| | | | block special files |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_BWRITE | | WRITE | VFS serializes reads from and writes to |
|
|
| | | | block special files |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_CHMOD | READ | WRITE | vmnt is only locked if file is not |
|
|
| | | | already opened |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_CHOWN | READ | WRITE | vmnt is only locked if file is not |
|
|
| | | | already opened |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_CREATE | WRITE | WRITE | The directory in which the file is |
|
|
| | | | created is write locked |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_FLUSH | | | Mutually exclusive to REQ_BREAD and |
|
|
| | | | REQ_BWRITE |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_FTRUNC | READ | WRITE | vmnt is only locked if file is not |
|
|
| | | | already opened |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_GETDENTS | READ | READ | vmnt is only locked if file is not |
|
|
| | | | already opened |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_INHIBREAD| | READ | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_LINK | READSER | WRITE | REQ_INODE_NR is locked READ |
|
|
| | | | REQ_DIR_INO is locked WRITE |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_LOOKUP | READSER | | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_MKDIR | READSER | WRITE | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_MKNOD | READSER | WRITE | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
|REQ_MOUNTPOINT| WRITE | WRITE | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
|REQ_NEW_DRIVER| | | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_NEWNODE | | | Only sent to PFS |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_PUTNODE | | READSER | READSER when dropping all but one |
|
|
| | | or WRITE| references. WRITE when final reference |
|
|
| | | | is dropped (i.e., no longer in use) |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_RDLINK | READ | READ | In some circumstances stricter locking |
|
|
| | | | might be applied, but not guaranteed |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_READ | | READ | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
|REQ_READSUPER | WRITE | | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_RENAME | WRITE | WRITE | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_RMDIR | WRITE | WRITE | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_SLINK | READSER | READ | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_STAT | READ | READ | vmnt is only locked if file is not |
|
|
| | | | already opened |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_STATVFS | READ | READ | vmnt is only locked if file is not |
|
|
| | | | already opened |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_SYNC | READ | | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_UNLINK | WRITE | WRITE | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_UNMOUNT | WRITE | | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_UTIME | READ | READ | |
|
|
+--------------+---------+---------+-----------------------------------------+
|
|
| REQ_WRITE | | WRITE | |
|
|
-----------------------------------------------------------------------------+
|
|
}}}
|
|
Table 7: VFS-FS requests locking guarantees
|
|
|
|
== Recovery from driver crashes ==
|
|
## 5 Recovery from driver crashes
|
|
VFS can recover from block special file and character special file driver
|
|
crashes. It can recover to some degree from a crashed File Server (which we
|
|
can regard as a driver).
|
|
|
|
=== Recovery from block drivers crashes ===
|
|
## 5.1 Recovery from block drivers crashes
|
|
When reading or writing, VFS doesn't communicate with block drivers directly,
|
|
but always through a File Server (the root File Server being default). If the
|
|
block driver crashes, the File Server does most of the work of the recovery
|
|
procedure. VFS loops through all open files for block special files that
|
|
were handled by this driver and reopens them. After that it sends the new
|
|
endpoint to the File Server so it can finish the recover procedure. Finally,
|
|
the File Server will retry pending requests if possible. However, reopening
|
|
files can cause the block driver to crash again. When that happens, VFS will
|
|
stop the recovery. A driver can return ERESTART to VFS to tell it to retry
|
|
a request. VFS does this with an arbitrary maximum of 5 attempts.
|
|
|
|
=== Recovery from character driver crashes ===
|
|
## 5.2 Recovery from character driver crashes
|
|
While VFS used to support minimal recovery from character driver crashes, the
|
|
added complexity has so far proven to outweigh the benefits, especially since
|
|
such crash recovery can never be fully transparent: it depends entirely on the
|
|
character device as to whether repeating an I/O request makes sense at all.
|
|
Currently, all operations except close(2) on a file descriptor that identifies
|
|
a device on a crashed character driver, will result in an EIO error. It is up
|
|
to the application to reopen the character device and retry whatever it was
|
|
doing in the appropriate manner. In the future, automatic reopen and I/O
|
|
restart may be reintroduced for a limited subset of character drivers.
|
|
|
|
=== Recovery from File Server crashes ===
|
|
## 5.3 Recovery from File Server crashes
|
|
At the time of writing we cannot recover from crashed File Servers. When
|
|
VFS detects it has to clean up the remnants of a File Server process (i.e.,
|
|
through an exit(2)), it marks all associated file descriptors as invalid
|
|
and cancels ongoing and pending requests to that File Server. Resources that
|
|
were in use by the File Server are cleaned up.
|
|
|
|
[0] http://wiki.minix3.org/en/DevelopersGuide/VfsFsProtocol
|
|
|
|
[1] http://www.cs.vu.nl/~dcvmoole/minix/blockchar.txt
|
|
|
|
[2] http://www.minix3.org/theses/moolenbroek-multimedia-support.pdf
|