linux

q3k/linux

History

Will Drewry e2cfabdfd0 seccomp: add system call filtering using BPF [This patch depends on luto@mit.edu's no_new_privs patch: https://lkml.org/lkml/2012/1/30/264 The whole series including Andrew's patches can be found here: https://github.com/redpig/linux/tree/seccomp Complete diff here: https://github.com/redpig/linux/compare/1dc65fed...seccomp ] This patch adds support for seccomp mode 2. Mode 2 introduces the ability for unprivileged processes to install system call filtering policy expressed in terms of a Berkeley Packet Filter (BPF) program. This program will be evaluated in the kernel for each system call the task makes and computes a result based on data in the format of struct seccomp_data. A filter program may be installed by calling: struct sock_fprog fprog = { ... }; ... prctl(PR_SET_SECCOMP, SECCOMP_MODE_FILTER, &fprog); The return value of the filter program determines if the system call is allowed to proceed or denied. If the first filter program installed allows prctl(2) calls, then the above call may be made repeatedly by a task to further reduce its access to the kernel. All attached programs must be evaluated before a system call will be allowed to proceed. Filter programs will be inherited across fork/clone and execve. However, if the task attaching the filter is unprivileged (!CAP_SYS_ADMIN) the no_new_privs bit will be set on the task. This ensures that unprivileged tasks cannot attach filters that affect privileged tasks (e.g., setuid binary). There are a number of benefits to this approach. A few of which are as follows: - BPF has been exposed to userland for a long time - BPF optimization (and JIT'ing) are well understood - Userland already knows its ABI: system call numbers and desired arguments - No time-of-check-time-of-use vulnerable data accesses are possible. - system call arguments are loaded on access only to minimize copying required for system call policy decisions. Mode 2 support is restricted to architectures that enable HAVE_ARCH_SECCOMP_FILTER. In this patch, the primary dependency is on syscall_get_arguments(). The full desired scope of this feature will add a few minor additional requirements expressed later in this series. Based on discussion, SECCOMP_RET_ERRNO and SECCOMP_RET_TRACE seem to be the desired additional functionality. No architectures are enabled in this patch. Signed-off-by: Will Drewry <wad@chromium.org> Acked-by: Serge Hallyn <serge.hallyn@canonical.com> Reviewed-by: Indan Zupancic <indan@nul.nu> Acked-by: Eric Paris <eparis@redhat.com> Reviewed-by: Kees Cook <keescook@chromium.org> v18: - rebase to v3.4-rc2 - s/chk/check/ (akpm@linux-foundation.org,jmorris@namei.org) - allocate with GFP_KERNEL\|__GFP_NOWARN (indan@nul.nu) - add a comment for get_u32 regarding endianness (akpm@) - fix other typos, style mistakes (akpm@) - added acked-by v17: - properly guard seccomp filter needed headers (leann@ubuntu.com) - tighten return mask to 0x7fff0000 v16: - no change v15: - add a 4 instr penalty when counting a path to account for seccomp_filter size (indan@nul.nu) - drop the max insns to 256KB (indan@nul.nu) - return ENOMEM if the max insns limit has been hit (indan@nul.nu) - move IP checks after args (indan@nul.nu) - drop !user_filter check (indan@nul.nu) - only allow explicit bpf codes (indan@nul.nu) - exit_code -> exit_sig v14: - put/get_seccomp_filter takes struct task_struct (indan@nul.nu,keescook@chromium.org) - adds seccomp_chk_filter and drops general bpf_run/chk_filter user - add seccomp_bpf_load for use by net/core/filter.c - lower max per-process/per-hierarchy: 1MB - moved nnp/capability check prior to allocation (all of the above: indan@nul.nu) v13: - rebase on to `88ebdda615` v12: - added a maximum instruction count per path (indan@nul.nu,oleg@redhat.com) - removed copy_seccomp (keescook@chromium.org,indan@nul.nu) - reworded the prctl_set_seccomp comment (indan@nul.nu) v11: - reorder struct seccomp_data to allow future args expansion (hpa@zytor.com) - style clean up, @compat dropped, compat_sock_fprog32 (indan@nul.nu) - do_exit(SIGSYS) (keescook@chromium.org, luto@mit.edu) - pare down Kconfig doc reference. - extra comment clean up v10: - seccomp_data has changed again to be more aesthetically pleasing (hpa@zytor.com) - calling convention is noted in a new u32 field using syscall_get_arch. This allows for cross-calling convention tasks to use seccomp filters. (hpa@zytor.com) - lots of clean up (thanks, Indan!) v9: - n/a v8: - use bpf_chk_filter, bpf_run_filter. update load_fns - Lots of fixes courtesy of indan@nul.nu: -- fix up load behavior, compat fixups, and merge alloc code, -- renamed pc and dropped __packed, use bool compat. -- Added a hidden CONFIG_SECCOMP_FILTER to synthesize non-arch dependencies v7: (massive overhaul thanks to Indan, others) - added CONFIG_HAVE_ARCH_SECCOMP_FILTER - merged into seccomp.c - minimal seccomp_filter.h - no config option (part of seccomp) - no new prctl - doesn't break seccomp on systems without asm/syscall.h (works but arg access always fails) - dropped seccomp_init_task, extra free functions, ... - dropped the no-asm/syscall.h code paths - merges with network sk_run_filter and sk_chk_filter v6: - fix memory leak on attach compat check failure - require no_new_privs \|\| CAP_SYS_ADMIN prior to filter installation. (luto@mit.edu) - s/seccomp_struct_/seccomp_/ for macros/functions (amwang@redhat.com) - cleaned up Kconfig (amwang@redhat.com) - on block, note if the call was compat (so the # means something) v5: - uses syscall_get_arguments (indan@nul.nu,oleg@redhat.com, mcgrathr@chromium.org) - uses union-based arg storage with hi/lo struct to handle endianness. Compromises between the two alternate proposals to minimize extra arg shuffling and account for endianness assuming userspace uses offsetof(). (mcgrathr@chromium.org, indan@nul.nu) - update Kconfig description - add include/seccomp_filter.h and add its installation - (naive) on-demand syscall argument loading - drop seccomp_t (eparis@redhat.com) v4: - adjusted prctl to make room for PR_[SG]ET_NO_NEW_PRIVS - now uses current->no_new_privs (luto@mit.edu,torvalds@linux-foundation.com) - assign names to seccomp modes (rdunlap@xenotime.net) - fix style issues (rdunlap@xenotime.net) - reworded Kconfig entry (rdunlap@xenotime.net) v3: - macros to inline (oleg@redhat.com) - init_task behavior fixed (oleg@redhat.com) - drop creator entry and extra NULL check (oleg@redhat.com) - alloc returns -EINVAL on bad sizing (serge.hallyn@canonical.com) - adds tentative use of "always_unprivileged" as per torvalds@linux-foundation.org and luto@mit.edu v2: - (patch 2 only) Signed-off-by: James Morris <james.l.morris@oracle.com>		2012-04-14 11:13:20 +10:00
..
debug	KGDB/KDB regression fixes	2012-04-04 17:26:08 -07:00
events	Merge branch 'linus' into perf/urgent	2012-03-26 17:19:03 +02:00
gcov	gcov: disable CONSTRUCTORS for UML	2011-07-26 16:49:45 -07:00
irq	Merge branch 'irq-core-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2012-03-30 18:08:05 -07:00
power	PM / QoS: add pm_qos_update_request_timeout() API	2012-03-28 23:31:24 +02:00
sched	Merge branch 'sched-urgent-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2012-03-31 13:35:31 -07:00
time	Merge branch 'timers-core-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2012-03-29 14:16:48 -07:00
trace	Merge branch 'akpm' (Andrew's patch-bomb)	2012-04-05 15:30:34 -07:00
.gitignore
acct.c	Merge branch 'for-linus2' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs	2012-01-08 12:19:57 -08:00
async.c	kernel/async: remove redundant declaration.	2012-01-13 09:32:18 +10:30
audit.c	constify path argument of audit_log_d_path()	2012-03-20 21:29:40 -04:00
audit.h	audit: remove AUDIT_SETUP_CONTEXT as it isn't used	2012-01-17 16:16:57 -05:00
audit_tree.c	audit_tree,rcu: Convert call_rcu(__put_tree) to kfree_rcu()	2011-07-20 14:10:11 -07:00
audit_watch.c
auditfilter.c	audit: allow interfield comparison in audit rules	2012-01-17 16:17:01 -05:00
auditsc.c	kernel-doc: fix new warnings in auditsc.c	2012-01-23 08:44:53 -08:00
backtracetest.c
bounds.c
capability.c	Revert "capabitlies: ns_capable can use the cap helpers rather than lsm call"	2012-01-17 10:19:41 -08:00
cgroup.c	cgroup: cgroup_attach_task() could return -errno after success	2012-03-29 22:03:33 -07:00
cgroup_freezer.c	cgroup: remove cgroup_subsys argument from callbacks	2012-02-02 09:20:22 -08:00
compat.c	compat: Add helper functions to read/write struct timeval, timespec	2012-02-20 12:48:47 -08:00
configs.c	kernel/configs.c: include MODULE_*() when CONFIG_IKCONFIG_PROC=n	2011-07-25 20:57:15 -07:00
cpu.c	Merge branch 'pm-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm	2012-01-08 13:10:57 -08:00
cpu_pm.c	cpu_pm: call notifiers during suspend	2011-09-23 12:05:29 +05:30
cpuset.c	Autogenerated GPG tag for Rusty D1ADB8F1: 15EE 8D6C AB0E 7F0C F999 BFCB D920 0E6C D1AD B8F1	2012-04-02 08:53:24 -07:00
crash_dump.c	Merge branch 'modsplit-Oct31_2011' of git://git.kernel.org/pub/scm/linux/kernel/git/paulg/linux	2011-11-06 19:44:47 -08:00
cred.c	security: trim security.h	2012-02-14 10:45:42 +11:00
delayacct.c	KVM: Steal time implementation	2011-07-14 12:59:14 +03:00
dma.c	Remove all #inclusions of asm/system.h	2012-03-28 18:30:03 +01:00
elfcore.c
exec_domain.c
exit.c	Merge branch 'x86-x32-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2012-03-29 18:12:23 -07:00
extable.c
fork.c	seccomp: add system call filtering using BPF	2012-04-14 11:13:20 +10:00
freezer.c	PM / Freezer: Remove references to TIF_FREEZE in comments	2012-03-04 23:08:54 +01:00
futex.c	futex: Mark get_robust_list as deprecated	2012-03-29 11:37:17 +02:00
futex_compat.c	futex: Mark get_robust_list as deprecated	2012-03-29 11:37:17 +02:00
groups.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
hrtimer.c	Merge branch 'timers-urgent-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2011-11-28 08:43:52 -08:00
hung_task.c	hung_task: fix the broken rcu_lock_break() logic	2012-03-05 15:49:42 -08:00
irq_work.c	irq_work: fix compile failure on MIPS from system.h split	2012-04-02 08:48:04 -07:00
itimer.c	[S390] cputime: add sparse checking and cleanup	2011-12-15 14:56:19 +01:00
jump_label.c	static keys: Inline the static_key_enabled() function	2012-02-28 20:01:08 +01:00
kallsyms.c
Kconfig.freezer
Kconfig.hz
Kconfig.locks	locking/kconfig: Simplify INLINE_SPIN_UNLOCK usage	2012-03-23 13:18:57 +01:00
Kconfig.preempt	locking/kconfig: Simplify INLINE_SPIN_UNLOCK usage	2012-03-23 13:18:57 +01:00
kexec.c	Merge branch 'akpm' (Andrew's patch-bomb)	2012-03-28 17:19:28 -07:00
kfifo.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
kmod.c	PM / Sleep: Mitigate race between the freezer and request_firmware()	2012-03-28 23:30:28 +02:00
kprobes.c	kprobes: return proper error code from register_kprobe()	2012-03-05 15:49:42 -08:00
ksysfs.c	kernel: ksysfs.c is implicitly using stat.h	2011-10-31 09:20:13 -04:00
kthread.c	freezer: kill unused set_freezable_with_signal()	2011-11-23 09:28:17 -08:00
latencytop.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
lockdep.c	lockdep: Add CPU-idle/offline warning to lockdep-RCU splat	2012-02-21 09:06:06 -08:00
lockdep_internals.h
lockdep_proc.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
lockdep_states.h
Makefile	sysctl: Move the implementation into fs/proc/proc_sysctl.c	2012-01-24 16:37:54 -08:00
module.c	module: Remove module size limit	2012-03-26 12:50:53 +10:30
mutex-debug.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
mutex-debug.h
mutex.c	sched/rt: Use schedule_preempt_disabled()	2012-03-01 10:28:03 +01:00
mutex.h
notifier.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
nsproxy.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
padata.c	padata: Fix cpu hotplug	2012-03-29 19:52:46 +08:00
panic.c	panic: don't print redundant backtraces on oops	2012-01-12 20:13:11 -08:00
params.c	params: <level>_initcall-like kernel parameters	2012-03-26 12:50:51 +10:30
pid.c	vfs: fix panic in __d_lookup() with high dentry hashtable counts	2012-02-13 20:45:38 -05:00
pid_namespace.c	pidns: add reboot_pid_ns() to handle the reboot syscall	2012-03-28 17:14:36 -07:00
posix-cpu-timers.c	[S390] cputime: add sparse checking and cleanup	2011-12-15 14:56:19 +01:00
posix-timers.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
printk.c	Merge branch 'sched-core-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2012-03-20 10:31:44 -07:00
profile.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
ptrace.c	ptrace: remove PTRACE_SEIZE_DEVEL bit	2012-03-23 16:58:41 -07:00
range.c	range: fix bogus misuse of module.h to get printk()	2011-10-31 09:20:11 -04:00
rcu.h	rcu: Allow nesting of rcu_idle_enter() and rcu_idle_exit()	2012-02-21 09:06:12 -08:00
rcupdate.c	rcu: Check for illegal use of RCU from offlined CPUs	2012-02-21 09:06:03 -08:00
rcutiny.c	rcu: Add RCU_NONIDLE() for idle-loop RCU read-side critical sections	2012-02-21 09:06:13 -08:00
rcutiny_plugin.h	rcu: Simplify unboosting checks	2012-02-21 09:03:43 -08:00
rcutorture.c	PTR_ERR should be called before its argument is cleared.	2012-02-21 09:06:10 -08:00
rcutree.c	rcu: Stop spurious warnings from synchronize_sched_expedited	2012-02-21 15:33:34 -08:00
rcutree.h	rcu: Rework detection of use of RCU by offline CPUs	2012-02-21 09:06:07 -08:00
rcutree_plugin.h	rcu: Hold off RCU_FAST_NO_HZ after timer posted	2012-02-21 09:42:30 -08:00
rcutree_trace.c	rcu: Rework detection of use of RCU by offline CPUs	2012-02-21 09:06:07 -08:00
relay.c	relay: prevent integer overflow in relay_open()	2012-02-10 09:04:49 +01:00
res_counter.c	net: introduce res_counter_charge_nofail() for socket allocations	2012-01-22 15:08:46 -05:00
resource.c	kernel/resource.c: move EXPORT_SYMBOL right after definition	2012-02-03 23:37:07 +01:00
rtmutex-debug.c	lockdep, rtmutex, bug: Show taint flags on error	2011-12-06 08:16:49 +01:00
rtmutex-debug.h
rtmutex-tester.c	rtmutex-tester: convert sysdev_class to a regular subsystem	2011-12-14 14:54:22 -08:00
rtmutex.c	Revert "rcu: Permit rt_mutex_unlock() with irqs disabled"	2011-12-11 10:33:18 -08:00
rtmutex.h
rtmutex_common.h
rwsem.c	Remove all #inclusions of asm/system.h	2012-03-28 18:30:03 +01:00
seccomp.c	seccomp: add system call filtering using BPF	2012-04-14 11:13:20 +10:00
semaphore.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
signal.c	Disintegrate and delete asm/system.h	2012-03-28 15:58:21 -07:00
smp.c	smp: add func to IPI cpus based on parameter func	2012-03-28 17:14:35 -07:00
softirq.c	Merge branch 'timers-core-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2012-03-20 10:32:09 -07:00
spinlock.c	locking/kconfig: Simplify INLINE_SPIN_UNLOCK usage	2012-03-23 13:18:57 +01:00
srcu.c	rcu: Call out dangers of expedited RCU primitives	2012-02-21 09:06:08 -08:00
stacktrace.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
stop_machine.c	Merge branch 'modsplit-Oct31_2011' of git://git.kernel.org/pub/scm/linux/kernel/git/paulg/linux	2011-11-06 19:44:47 -08:00
sys.c	seccomp: add system call filtering using BPF	2012-04-14 11:13:20 +10:00
sys_ni.c	Cross Memory Attach	2011-10-31 17:30:44 -07:00
sysctl.c	sysctl: fix write access to dmesg_restrict/kptr_restrict	2012-04-05 14:51:43 +10:00
sysctl_binary.c	binary_sysctl(): fix memory leak	2011-12-20 10:25:04 -08:00
taskstats.c	Make TASKSTATS require root access	2011-09-19 17:04:37 -07:00
test_kprobes.c
time.c	time: Remove bogus comments	2012-03-15 18:17:55 -07:00
timeconst.pl
timer.c	Merge branch 'core-debugobjects-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2012-01-06 07:53:34 -08:00
tracepoint.c	static keys: Introduce 'struct static_key', static_key_true()/false() and static_key_slow_[inc\|dec]()	2012-02-24 10:05:59 +01:00
tsacct.c	[S390] cputime: add sparse checking and cleanup	2011-12-15 14:56:19 +01:00
uid16.c
up.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
user-return-notifier.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
user.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
user_namespace.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
utsname.c	kernel: Map most files to use export.h instead of module.h	2011-10-31 09:20:12 -04:00
utsname_sysctl.c	Merge branch 'modsplit-Oct31_2011' of git://git.kernel.org/pub/scm/linux/kernel/git/paulg/linux	2011-11-06 19:44:47 -08:00
wait.c	lockdep/waitqueues: Add better annotation	2011-12-21 10:07:39 +01:00
watchdog.c	kernel/watchdog.c: add comment to watchdog() exit path	2012-03-23 16:58:32 -07:00
workqueue.c	Merge branch 'for-3.4' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/wq	2012-03-20 18:13:22 -07:00
workqueue_sched.h