提交 · 3a053b1a30dcb4e39569bcce2f4357509260db75 · openeuler / Kernel

02 3月, 2018 1 次提交

net: Fix spelling mistake "greater then" -> "greater than" · 3a053b1a

由 Gal Pressman 提交于 2月 28, 2018

Fix trivial spelling mistake "greater then" -> "greater than".
Signed-off-by: NGal Pressman <galp@mellanox.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

3a053b1a

01 3月, 2018 1 次提交

net: fib_rules: support for match on ip_proto, sport and dport · bfff4862

由 Roopa Prabhu 提交于 2月 28, 2018

uapi for ip_proto, sport and dport range match
in fib rules.
Signed-off-by: NRoopa Prabhu <roopa@cumulusnetworks.com>
Acked-by: NNikolay Aleksandrov <nikolay@cumulusnetworks.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

bfff4862

27 2月, 2018 1 次提交

net: make kmem caches as __ro_after_init · 08009a76

由 Alexey Dobriyan 提交于 2月 24, 2018

All kmem caches aren't reallocated once set up.
Signed-off-by: NAlexey Dobriyan <adobriyan@gmail.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

08009a76

24 2月, 2018 2 次提交

net: fib_rules: Add new attribute to set protocol · 1b71af60

由 Donald Sharp 提交于 2月 23, 2018

For ages iproute2 has used `struct rtmsg` as the ancillary header for
FIB rules and in the process set the protocol value to RTPROT_BOOT.
Until ca56209a66 ("net: Allow a rule to track originating protocol")
the kernel rules code ignored the protocol value sent from userspace
and always returned 0 in notifications. To avoid incompatibility with
existing iproute2, send the protocol as a new attribute.

Fixes: cac56209 ("net: Allow a rule to track originating protocol")
Signed-off-by: NDonald Sharp <sharpd@cumulusnetworks.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

1b71af60

net_sched: gen_estimator: fix broken estimators based on percpu stats · a5f7add3

由 Eric Dumazet 提交于 2月 22, 2018

pfifo_fast got percpu stats lately, uncovering a bug I introduced last
year in linux-4.10.

I missed the fact that we have to clear our temporary storage
before calling __gnet_stats_copy_basic() in the case of percpu stats.

Without this fix, rate estimators (tc qd replace dev xxx root est 1sec
4sec pfifo_fast) are utterly broken.

Fixes: 1c0d32fd ("net_sched: gen_estimator: complete rewrite of rate estimators")
Signed-off-by: NEric Dumazet <edumazet@google.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

a5f7add3

22 2月, 2018 3 次提交

bpf: clean up unused-variable warning · a7dcdf6e

由 Arnd Bergmann 提交于 2月 20, 2018

The only user of this variable is inside of an #ifdef, causing
a warning without CONFIG_INET:

net/core/filter.c: In function '____bpf_sock_ops_cb_flags_set':
net/core/filter.c:3382:6: error: unused variable 'val' [-Werror=unused-variable]
  int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;

This replaces the #ifdef with a nicer IS_ENABLED() check that
makes the code more readable and avoids the warning.

Fixes: b13d8807 ("bpf: Adds field bpf_sock_ops_cb_flags to tcp_sock")
Signed-off-by: NArnd Bergmann <arnd@arndb.de>
Signed-off-by: NDaniel Borkmann <daniel@iogearbox.net>

a7dcdf6e

net: Allow a rule to track originating protocol · cac56209

由 Donald Sharp 提交于 2月 20, 2018

Allow a rule that is being added/deleted/modified or
dumped to contain the originating protocol's id.

The protocol is handled just like a routes originating
protocol is.  This is especially useful because there
is starting to be a plethora of different user space
programs adding rules.

Allow the vrf device to specify that the kernel is the originator
of the rule created for this device.
Signed-off-by: NDonald Sharp <sharpd@cumulusnetworks.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

cac56209

tcp: switch to GSO being always on · 0a6b2a1d

由 Eric Dumazet 提交于 2月 19, 2018

Oleksandr Natalenko reported performance issues with BBR without FQ
packet scheduler that were root caused to lack of SG and GSO/TSO on
his configuration.

In this mode, TCP internal pacing has to setup a high resolution timer
for each MSS sent.

We could implement in TCP a strategy similar to the one adopted
in commit fefa569a ("net_sched: sch_fq: account for schedule/timers drifts")
or decide to finally switch TCP stack to a GSO only mode.

This has many benefits :

1) Most TCP developments are done with TSO in mind.
2) Less high-resolution timers needs to be armed for TCP-pacing
3) GSO can benefit of xmit_more hint
4) Receiver GRO is more effective (as if TSO was used for real on sender)
   -> Lower ACK traffic
5) Write queues have less overhead (one skb holds about 64KB of payload)
6) SACK coalescing just works.
7) rtx rb-tree contains less packets, SACK is cheaper.

This patch implements the minimum patch, but we can remove some legacy
code as follow ups.

Tested:

On 40Gbit link, one netperf -t TCP_STREAM

BBR+fq:
sg on:  26 Gbits/sec
sg off: 15.7 Gbits/sec   (was 2.3 Gbit before patch)

BBR+pfifo_fast:
sg on:  24.2 Gbits/sec
sg off: 14.9 Gbits/sec  (was 0.66 Gbit before patch !!! )

BBR+fq_codel:
sg on:  24.4 Gbits/sec
sg off: 15 Gbits/sec  (was 0.66 Gbit before patch !!! )
Signed-off-by: NEric Dumazet <edumazet@google.com>
Reported-by: NOleksandr Natalenko <oleksandr@natalenko.name>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

0a6b2a1d

21 2月, 2018 4 次提交

devlink: Move size validation to core · cc944ead

由 Arkadi Sharshevsky 提交于 2月 20, 2018

Currently the size validation is done via a cb, which is unneeded. The
size validation can be moved to core. The next patch will perform cleanup.
Signed-off-by: NArkadi Sharshevsky <arkadis@mellanox.com>
Signed-off-by: NJiri Pirko <jiri@mellanox.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

cc944ead

net: Queue net_cleanup_work only if there is first net added · 8349efd9

由 Kirill Tkhai 提交于 2月 19, 2018

When llist_add() returns false, cleanup_net() hasn't made its
llist_del_all(), while the work has already been scheduled
by the first queuer. So, we may skip queue_work() in this case.
Signed-off-by: NKirill Tkhai <ktkhai@virtuozzo.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

8349efd9

net: Make cleanup_list and net::cleanup_list of llist type · 65b7b5b9

由 Kirill Tkhai 提交于 2月 19, 2018

This simplifies cleanup queueing and makes cleanup lists
to use llist primitives. Since llist has its own cmpxchg()
ordering, cleanup_list_lock is not more need.

Also, struct llist_node is smaller, than struct list_head,
so we save some bytes in struct net with this patch.
Signed-off-by: NKirill Tkhai <ktkhai@virtuozzo.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

65b7b5b9

net: Kill net_mutex · 19efbd93

由 Kirill Tkhai 提交于 2月 19, 2018

We take net_mutex, when there are !async pernet_operations
registered, and read locking of net_sem is not enough. But
we may get rid of taking the mutex, and just change the logic
to write lock net_sem in such cases. This obviously reduces
the number of lock operations, we do.
Signed-off-by: NKirill Tkhai <ktkhai@virtuozzo.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

19efbd93

17 2月, 2018 2 次提交

sock: permit SO_ZEROCOPY on PF_RDS socket · 28190752

由 Sowmini Varadhan 提交于 2月 15, 2018

allow the application to set SO_ZEROCOPY on the underlying sk
of a PF_RDS socket
Signed-off-by: NSowmini Varadhan <sowmini.varadhan@oracle.com>
Acked-by: NWillem de Bruijn <willemb@google.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

28190752

skbuff: export mm_[un]account_pinned_pages for other modules · 6f89dbce

由 Sowmini Varadhan 提交于 2月 15, 2018

RDS would like to use the helper functions for managing pinned pages
added by Commit a91dbff5 ("sock: ulimit on MSG_ZEROCOPY pages")
Signed-off-by: NSowmini Varadhan <sowmini.varadhan@oracle.com>
Acked-by: NWillem de Bruijn <willemb@google.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

6f89dbce

15 2月, 2018 2 次提交

net: fix race on decreasing number of TX queues · ac5b7019

由 Jakub Kicinski 提交于 2月 12, 2018

netif_set_real_num_tx_queues() can be called when netdev is up.
That usually happens when user requests change of number of
channels/rings with ethtool -L.  The procedure for changing
the number of queues involves resetting the qdiscs and setting
dev->num_tx_queues to the new value.  When the new value is
lower than the old one, extra care has to be taken to ensure
ordering of accesses to the number of queues vs qdisc reset.

Currently the queues are reset before new dev->num_tx_queues
is assigned, leaving a window of time where packets can be
enqueued onto the queues going down, leading to a likely
crash in the drivers, since most drivers don't check if TX
skbs are assigned to an active queue.

Fixes: e6484930 ("net: allocate tx queues in register_netdevice")
Signed-off-by: NJakub Kicinski <jakub.kicinski@netronome.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

ac5b7019

net: Make dn_ptr depend on CONFIG_DECNET · 330c7272

由 David Ahern 提交于 2月 13, 2018

Signed-off-by: NDavid Ahern <dsahern@gmail.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

330c7272

13 2月, 2018 16 次提交

net: Convert diag_net_ops · 59a51358